Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.
Figure 2: Slides shortcuts comparison between SlideLab (Ours) and DeepPresenter, Kimi Slides, Manus.
Per slide
Whole talk
System
Grounding. errors ↓
Fig. errors ↓
Design ↑
Figure use ↑
Narrative. errors ↓
Coverage ↑
Env. score ↑
PPTAgent
0.98
0.64
3.21
3.18
5.40
0.61
0.42
DeepPresenter
2.08
0.53
3.49
3.27
5.19
0.94
0.32
Kimi Slides
1.48
0.17
3.67
3.08
4.56
0.91
0.58
Manus
1.81
–
3.59
–
4.14
0.89
0.37
SlideLab (ours)
0.71
0.26
4.21
3.80
3.29
0.99
0.73
Table 1: ConfArena results on 100 papers, with three retained decks per paper for SlideLab, DeepPresenter, and PPTAgent, and one for Kimi Slides and Manus. Best values are bold.
System
Papers picked best
% of 30
SlideLab (ours)
23
77%
Kimi Slides
7
23%
DeepPresenter
0
0%
Manus
0
0%
Table 2: Results of the blind human preference study on 30 research papers. Annotators selected the presentation they would use to deliver the paper at a conference after evaluating factual correctness, coverage, narrative flow, visual design, and figure usage.
Configuration
Env. score
SlideLab (full)
0.70
− Planner
0.51
− LayoutDebugger
0.58
− Compositor
0.65
− Custom visuals (paper figs only)
0.63
Table 3: Component ablations of SlideLab. The Planner and LayoutDebugger are the two stages whose removal changes the result significantly. Each configuration is run once per paper on the 100 -paper set, unlike Table 1 , which averages three independent runs per paper.
PPTEval (1–5)
PresentBench
SlidesGen-Bench
System
Content ↑
Design ↑
Coherence ↑
pass % ↑
quiz acc. ↑
PPTAgent
3.5
3.8
3.6
63
0.74
DeepPresenter
3.7
3.7
3.4
64
0.76
Kimi Slides
3.9
4.2
4.0
73
0.79
Manus
3.6
3.8
3.5
61
0.72
SlideLab (ours)
4.1
4.4
4.2
71
0.84
Table 4: Evaluations on the 100 -paper set using other three methods. PPTEval: LLM-judge ratings of content, design, and coherence ( 1 – 5 ). PresentBench: fraction of decks passing the benchmark’s checks (%). SlidesGenBench: accuracy on quizzes the benchmark generates from the source paper. ↑ higher is better. Best in bold.
System
Human
ConfArena
PPTEval
PresentBench
SlidesGenB.
SlideLab
1
1
1
2
1
Kimi Slides
2
2
2
1
2
Manus
3
3
3
4
4
DeepPresenter
3
4
4
3
3
Table 5: Comparison with human preferences. (a) Aggregate system rankings ( 1 = best). Manus and DeepPresenter tie in the human study; PPTEval ranks use the mean of its three dimensions. (b) Frequency with which ConfArena selects the same preferred deck as the human majority for a paper or as an individual annotator.
Perturbation
PresentB.
PPTEval
SlidesG.
ConfArena
Falsified num.
×
✓
×
✓
Degraded fig.
✓
✓
✓
✓
Dropped slide
×
∘
✓
✓
Shuffled order
✓
✓
✓
✓
Caught /4
2
3
3
4
Table 6: Can evaluation frameworks detect planted errors in slides? ✓ = the framework’s relevant metric moved in the expected direction; ×= did not; ∘= the framework has no metric targeting that failure.
Mean change (positive = worse)
Damage applied
Arc defects
Grounding errors
Figure errors
Shuffle slide order
+7.25
+0.11
-0.01
Inject one false number
+0.50
+1.51
+0.20
Shrink figures by 50%
+0.25
-0.17
+0.51
Drop Slides
+2.50
+0.01
-0.01
Table 7: We damage 15 SlideLab decks in targeted ways and re-run the ConfArena. Cells show the mean change per metric; bold marks the metric the damage targets. Each damage moves mainly its own metric.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Sub-stage
Model
Max rounds / calls
Planner
Candidate plan generation
GPT-5.5
60
Candidate critic and selector
GLM-5.2
1
Plan grounding verifier
GPT-5.5
1
Slide Generator
Slide generation
MiMo-v2.5-Pro
50
Visual Generator
Visual instruction / prompt generation
MiMo-v2.5-Pro
15
Image generation
GPT-Image-2
1 per generated visual
Appendix
Table 8: Models and maximum round each SlideLab stage.
ConfArena score ↑
Configuration
GPT-5.4
Gemini-3.8- Flash
SlideLab (full) (MiMo-V2.5-Pro)
0.74
0.78
− Planner
0.53
0.62
− LayoutDebugger
0.60
0.69
− Compositor
0.70
0.74
− Custom visuals (paper figures only)
0.69
0.70
Appendix
Table 9: Component ablations on 30 papers using MiMo-V2.5-Pro for all text agents. Columns report environment scores under two ConfArena runs with different model. Each minus denotes removal of one component from the full pipeline.
Per slide
Whole talk
System
Grounding errors ↓
Figure errors ↓
Design ↑
Figure use ↑
Narrative errors ↓
Coverage ↑
Env. score ↑
PPTAgent (GPT-5.5)
1.462
0.443
3.441
3.090
5.13
0.64
0.38
DeepPresenter (GPT-5.5)
1.882
0.316
3.725
3.368
4.77
0.753
0.30
Kimi Slides
1.162
0.156
4.430
2.830
3.94
0.924
0.61
Manus
1.650
–
3.990
–
3.20
0.691
0.31
SlideLab (ours) (MiMo-V2.5-Pro, all text stages)
1.244
0.140
4.673
4.525
2.400
0.927
0.784
Appendix
Table 10: System comparison on 30 human-annotated papers using Gemini-3.8-Flash for all ConfArena evaluation roles. SlideLab uses MiMo-V2.5-Pro for all text agents. Metric definitions follow Table 1 . Bold marks the best value in each column; dashes indicate unavailable figure metrics.
Figure 3: ConfArena evaluation architecture: sequential slide review by the examiner and three independent attendees, audience Q&A and question diagnosis, and aggregation of per-slide and whole-talk evidence into the environment score.
Attendee
Focus and Question Trigger
Expert reviewer
Focuses on scientific rigor and correctness. Asks a question when a claim, result, or numerical value appears unsupported, inconsistent, or overstated.
Learner
Focuses on whether the core ideas, methodology, and motivation are easy to follow. Asks a question when an important concept or step is insufficiently explained.
Cross-field attendee
Focuses on accessibility without deep subfield knowledge. Asks a question when jargon or field-specific assumptions are introduced without adequate explanation.
Appendix
Table 11: Attendee personas used in ConfArena.
Metric
Pearson correlation ( r )
Grounding errors
0.939
Figure errors
0.952
Design
0.981
Figure use
0.917
Coverage
0.637
Narrative errors
0.552
Appendix
Table 12: Correlation between ConfArena scores and scores from an examiner that evaluates slides separately. Both use Gemini-3.8-Flash.
Narrative errors ↓
Evaluator
SlideLab
Kimi Slides
Static examiner
0.90
0.80
ConfArena
2.40
3.94
Appendix
Table 13: Mean narrative errors per deck under the static examiner and ConfArena.
Evaluator condition
Narrative errors per deck
Full ConfArena
3.88
Neutral attendee roles
3.40
Without Q&A
2.30
Examiner only
1.75
Appendix
Table 14: Mean narrative errors per deck under different ConfArena configurations.
Stage
Time (s)
Cost ($)
Tokens (M)
Planner
96
0.10
0.09
Slide Generator
264
0.03
0.23
ImageVisualGenerator + Compositor
261
0.42
0.32
LayoutDebugger
122
0.03
0.05
SlideLab total
743
0.58
0.69
DeepPresenter
1620
2.1
2.9
Appendix
Table 15: Average wall-clock time (end to end), API cost, and token usage per deck.
Figure 4: Annotation interface for the human preference study. Annotators see the four anonymized decks side by side, pick the one they would present, then rate it on six dimensions.
Dimension
What annotators are asked
Content Grounding
Do the slides match what the paper says?
Content Coverage
Does the deck cover the paper’s key contributions?
Narrative Structure
Do the slides tell a coherent story?
Visual Design
Does it look like a professional conference talk?
Information Density
Are slides concise, not walls of text?
Figure Usage
Are figures well-chosen and sized correctly?
Appendix
Table 16: The six dimensions annotators rate after picking their preferred deck. Each uses a five-point scale with the anchors shown in the interface.
Perturbation
Benchmark
Metric (direction)
Baseline
Perturbed
Δ
Detected?
Falsified number
PPTEval
content (1–5, ↑ )
4.30
4.27
−0.03
✓
SlidesGen-Bench
quiz accuracy (0–1, ↑ )
0.98
0.98
+0.00
×
PresentBench
material-dep. pass % ( ↑ )
27.5
45.0
+17.5
×
ConfArena
faithfulness (0–1, ↑ )
0.58
0.57
−0.01
✓
Degraded figure
PPTEval
design (1–5, ↑ )
4.05
3.85
−0.20
✓
SlidesGen-Bench
visual weighted total (0–10, ↑ )
8.45
8.22
−0.23
✓
Appendix
Table 17: Mean baseline and perturbed scores per benchmark per perturbation, across 15 papers. Δ = perturbed minus baseline; for metrics where higher is better, a negative Δ means the benchmark moved in the expected direction. ✓ = detected, ×= not detected, ∘= no metric exists for that failure mode. Baselines are the mean over the 15 unperturbed decks.
Figure 5: Representative qualitative comparisons supporting the quantitative results reported in the main paper.