PPaperPicks

Timothy Baldwin

Mohamed bin Zayed University of Artificial Intelligence, UAE

40 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Multilingual Social Bias Benchmark Incorporating Thinking Processes
  2. Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
  3. COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
  4. Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
  5. Do Diacritics Matter? Evaluating the Impact of Arabic Diacritics on Tokenization and LLM Benchmarks
  6. Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
  7. Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation
  8. On the Interplay between Human Label Variation and Model Fairness
  9. SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
  10. ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning
  11. Uncertainty Quantification for Large Language Models
  12. A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
  13. An Ethical Dataset from Real-World Interactions Between Users and Large Language Models
  14. Arabic Dataset for LLM Safeguard Evaluation
  15. Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
  16. BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities
  17. Bits Leaked per Query: Information-Theoretic Bounds for Adversarial Attacks on LLMs
  18. Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
  19. Evaluating Evidence Attribution in Generated Fact Checking Explanations
  20. Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models
  21. Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection Capabilities
  22. Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
  23. NAT: Enhancing Agent Tuning with Negative Samples
  24. Palo: A Polyglot Large Multimodal Model for 5B People
  25. Qorǵau: Evaluating Safety in Kazakh-Russian Bilingual Contexts
  26. Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
  27. ToolGen: Unified Tool Retrieval and Calling via Generation
  28. Uncertainty Quantification for Large Language Models
  29. Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
  30. A Chinese Dataset for Evaluating the Safeguards in Large Language Models
  31. ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
  32. Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
  33. BiMediX: Bilingual Medical Mixture of Experts LLM
  34. CMMLU: Measuring massive multitask language understanding in Chinese
  35. Demystifying Instruction Mixing for Fine-tuning Large Language Models
  36. Emergent Word Order Universals from Cognitively-Motivated Language Models
  37. Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
  38. Psychometric Predictive Power of Large Language Models
  39. Revisiting subword tokenization: A case study on affixal negation in large language models
  40. Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs