Large Language Model Benchmarks

Latest papers 184

All topics
CardsList
  1. FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

    May 23, 2025Haoyu Sun, Huichen Will Wang, Jiawei Gu +2Large Language Model Benchmarks

  2. BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity?

    Mar 19, 2025Pierre Chambon, Baptiste Roziere, Benoit Sagot +1Code GenerationLarge Language Model Benchmarks

  3. Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

    Mar 9, 2025Batu Guan, Xiao Wu, Yuanyuan Yuan +1Large Language Model BenchmarksReasoning Benchmark

  4. Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

    Jan 20, 2025Zijin Hong, Hao Wu, Su Dong +8Mathematical ReasoningLarge Language Model Benchmarks