PPaperPicks

José Hernández-Orallo

13 papers at tracked venues · 6 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. TRACE: A Corpus of Team Creative Discussions
  2. Contamination Budget: Trade-offs Between Breadth, Depth and Difficulty
  3. Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
  4. LLM GameLab: An Interactive Platform for Testing Large Language Models in Board Games
  5. Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
  6. Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
  7. PredictaBoard: Benchmarking LLM Score Predictability
  8. Relative Drawing Identification Complexity Is Invariant to Modality in Vision-Language Models
  9. Caveats and Solutions for Characterising General-Purpose AI
    ECAI 2024 · José Hernández-Orallo
  10. Distilling the Effects of Language Model Contamination
  11. Language Task Difficulty Prediction Through LLM-Annotated Meta-Features
  12. Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence
  13. Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint)