Language Model Ensembles

Momentum

15 papers in the last four weeks, up 400% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 93

All topics
CardsList
  1. TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

    Jul 22, 2026Isabel Xu, Cynthia Xu, Rachel Ren +2Multi-Agent LLM SystemsFinancial Sentiment Analysis

  2. Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift

    Jul 19, 2026Junade AliEnsemble LearningLanguage Model Ensembles

  3. Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing

    Jul 13, 2026Amal Saad Alshehri, Nelly Bencomo, Amir Atapour-AbarghoueiLegal NLPLegal IR

  4. Collective Intelligence with Foundation Models

    Jul 6, 2026J. de Curtò, I. de ZarzàAI Agent EvaluationMulti-Agent Collaboration

  5. Decentralized Aggregation of LLM Predictions via Wagering Mechanisms

    Jul 5, 2026Yuhong Luo, David M. Pennock, Xintong WangLanguage Model EnsemblesPrediction Markets

  6. LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution

    Jul 1, 2026Zhao Tian, Yingquan Zhao, Chenyao Suo +2LLM EvaluationAutomated Program Repair

  7. Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models

    Jun 30, 2026Justin Brenne, Christian MeskeLLM Uncertainty EstimationLanguage Model Ensembles

  8. Diversity is the Strength of the AI Crowd

    Jun 29, 2026Matthew Aitchison, Scott Jeen, Toby Shevlane +1Forecasting BenchmarksLanguage Model Ensembles

  9. Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch

    Jun 27, 2026Katja Berčič, Slobodan StanojevikjLLM-as-a-JudgeClassification

  10. The Capability Frontier: Benchmarks Miss 82% of Model Performance

    Jun 25, 2026Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8LLM EvaluationLanguage Model Generation Evaluation

  11. Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models

    Jun 24, 2026Mayumi Mohan, Ju-Hung Chen, Alexis E. BlockEvidence SelectionSmall Language Models

  12. Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

    Jun 3, 2026Ahmed Alansary, Molham Mohamed, Ali HamdiMedical QACurriculum Learning

  13. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

    Jun 3, 2026Ahmed Alansary, Ali HamdiText SummarizationLanguage Model Ensembles

  14. DLLG: Dynamic Logit-Level Gating of LLM Experts

    Jun 3, 2026Bingnan Li, Zhaoyang Zhang, Xiaoze Liu +6Ensemble LearningLanguage Model Ensembles

  15. q0: Primitives for Hyper-Epoch Pretraining

    Jun 2, 2026Bishwas Mandal, Shmuel Berman, Akshay Vegesna +1Language Model PretrainingEfficient Language Model Training

  16. CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

    Jun 2, 2026Alexander Apartsin, Yehudit ApersteinModel SelectionLLM Evaluation

  17. A Finite-Calibration Regime Map for LLM Judge Panels

    May 31, 2026Bin Zhu, Yanghui RaoLLM-as-a-JudgeLanguage Model Calibration

  18. Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

    May 29, 2026Shuheng Cao, Ruiqi Chen, Renjie Cao +3Named Entity RecognitionLLM-Assisted Annotation

  19. Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

    May 28, 2026Zhihao Wu, Gracia Gong, Qinglin Zhu +2Watermark RemovalLLM Watermarking