Generative AI Evaluation

Latest papers 211

All topics
CardsList
  1. Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task

    Apr 28, 2026Thomas A. Grossman, Yuan Chen, Sopiko DatuashviliML ReproducibilitySpreadsheet Automation

  2. Geographic Bias and Diversity in AI Evaluation

    Apr 28, 2026Zilong Liu, Krzysztof Janowicz, Gengchen Mai +2Algorithmic BiasLanguage Model Bias Evaluation

  3. MetaGAI: A Large-Scale and High-Quality Benchmark for Generative AI Model and Data Card Generation

    Apr 26, 2026Haoxuan Zhang, Ruochi Li, Yang Zhang +4Automated EvaluationGenerative AI Evaluation

  4. Ethics Testing: Proactive Identification of Generative AI System Harms

    Apr 23, 2026Shin Hwei Tan, Haibo Wang, Heng LiAI Safety EvaluationAutomated Test Generation

  5. Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

    Apr 22, 2026Rebecca L. JohnsonAI GovernanceGenerative AI Evaluation

  6. What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review

    Apr 21, 2026Ming JinAutomated Peer ReviewGenerative AI Evaluation

  7. Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

    Apr 20, 2026Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández +3Retrieval-Augmented GenerationEvidence-Grounded Generation

  8. SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

    Apr 16, 2026Dapeng Wu, Shun Lei, Wei Tan +5Text-to-Music GenerationMusic Generation Evaluation

  9. Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literature

    Apr 14, 2026Jinkai Tao, Yubo Wang, Xiaoyu Liu +1Scientific DiscoveryScientific Hypothesis Generation

  10. Artificial Intelligence Index Report 2026

    Apr 14, 2026Sha Sajadieh, Loredana Fattorini, Raymond Perrault +20AI for ScienceAI Governance

  11. When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

    Mar 27, 2026Juan Gabriel Kostelec, Qinghai GuoLLM Inference EfficiencyLLM Evaluation

  12. Overreliance on AI in Information-seeking from Video Content

    Mar 20, 2026Anders Giovanni Møller, Elisa Bassignana, Francesco Pierri +1Video QATrust in AI

  13. Adaptive Contracts for Cost-Effective AI Delegation

    Mar 17, 2026Eden Saig, Tamar Garbuz, Ariel D. Procaccia +2Generative AI Evaluation

  14. Quantifying the Effect of Test Set Contamination on Generative Evaluations

    Jan 7, 2026Rylan Schaeffer, Joshua Kazdan, Baber Abbasi +8Language Model Generation EvaluationBenchmark Contamination

  15. Developing an LLM-Based Feedback System Grounded in Evidence-Centered Design to Support Physics Problem Solving

    Dec 11, 2025Holger Maus, Fabian Kieser, Stefan Petersen +2Educational TechnologyGenerative AI in Education

  16. Generalized Design Choices for Deepfake Detectors

    Nov 26, 2025Lorenzo Pellegrini, Serafino Pandolfini, Davide Maltoni +3Deep Learning OptimizationDeepfake Detection

  17. Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot

    Nov 12, 2025Thomas D. Hull, Lizhe Zhang, Caitlin A. Stamatis +2HealthcareMental Health

  18. AtomBench: A Benchmarking Framework for Generative Crystal Reconstruction Models in Conventional Superconductors

    Oct 17, 2025Charles Rhys Campbell, Aldo H. Romero, Kamal ChoudharyMaterials ScienceGenerative Modeling

  19. Generative AI and Sales Productivity: Field Experiments in Online Retail

    Oct 14, 2025Lu Fang, Zhe Yuan, Kaifu Zhang +2Generative AI Evaluation

  20. Generative AI performance in core undergraduate mathematics: a curriculum-level case study

    Sep 15, 2025Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda +1Educational AssessmentGenerative AI in Education

  21. Towards AI-Assisted Research Writing: Benchmarking LLMs for AI/ML Introduction Generation

    Aug 19, 2025Krishna Garg, Firoz Shaik, Sambaran Bandyopadhyay +1LLM EvaluationLanguage Model Generation Evaluation

  22. Evaluating Style-Personalized Text Generation: Challenges and Directions

    Aug 8, 2025Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan +2LLM-as-a-JudgeLanguage Model Generation Evaluation

  23. Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

    May 27, 2024Xuemei Gu, Mario KrennAI-Assisted Scientific ResearchScientific Hypothesis Generation

  24. PQMass: Probabilistic Assessment of the Quality of Generative Models using Probability Mass Estimation

    Feb 6, 2024Pablo Lemos, Sammy Sharief, Esmeralda S. Whitammer +4Two-Sample TestingGenerative AI Evaluation

  25. LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics

    Date pendingAnselm Strohmaier, Samira Bödefeld, Oliver Straser +1Generative AI in EducationGenerative AI Evaluation