SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
Organizations: INSAIT, Sofia University “St. Kliment Ohridski”
Abstract
Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.
Figures & tables
| Per slide | Whole talk | ||||||
| System | Grounding. errors | Fig. errors | Design | Figure use | Narrative. errors | Coverage | Env. score |
| PPTAgent | 0.98 | 0.64 | 3.21 | 3.18 | 5.40 | 0.61 | 0.42 |
| DeepPresenter | 2.08 | 0.53 | 3.49 | 3.27 | 5.19 | 0.94 | 0.32 |
| Kimi Slides | 1.48 | 0.17 | 3.67 | 3.08 | 4.56 | 0.91 | 0.58 |
| Manus | 1.81 | – | 3.59 | – | 4.14 | 0.89 | 0.37 |
| SlideLab (ours) | 0.71 | 0.26 | 4.21 | 3.80 | 3.29 | 0.99 | 0.73 |
| System | Papers picked best | % of 30 |
| SlideLab (ours) | 23 | 77% |
| Kimi Slides | 7 | 23% |
| DeepPresenter | 0 | 0% |
| Manus | 0 | 0% |
| Configuration | Env. score |
| SlideLab (full) | 0.70 |
| Planner | 0.51 |
| LayoutDebugger | 0.58 |
| Compositor | 0.65 |
| Custom visuals (paper figs only) | 0.63 |
| PPTEval (1–5) | PresentBench | SlidesGen-Bench | |||
| System | Content | Design | Coherence | pass % | quiz acc. |
| PPTAgent | 3.5 | 3.8 | 3.6 | 63 | 0.74 |
| DeepPresenter | 3.7 | 3.7 | 3.4 | 64 | 0.76 |
| Kimi Slides | 3.9 | 4.2 | 4.0 | 73 | 0.79 |
| Manus | 3.6 | 3.8 | 3.5 | 61 | 0.72 |
| SlideLab (ours) | 4.1 | 4.4 | 4.2 | 71 | 0.84 |
| System | Human | ConfArena | PPTEval | PresentBench | SlidesGenB. |
| SlideLab | 1 | 1 | 1 | 2 | 1 |
| Kimi Slides | 2 | 2 | 2 | 1 | 2 |
| Manus | 3 | 3 | 3 | 4 | 4 |
| DeepPresenter | 3 | 4 | 4 | 3 | 3 |
| Perturbation | PresentB. | PPTEval | SlidesG. | ConfArena |
| Falsified num. | ✓ | ✓ | ||
| Degraded fig. | ✓ | ✓ | ✓ | ✓ |
| Dropped slide | ✓ | ✓ | ||
| Shuffled order | ✓ | ✓ | ✓ | ✓ |
| Caught /4 | 2 | 3 | 3 | 4 |
| Mean change (positive worse) | |||
| Damage applied | Arc defects | Grounding errors | Figure errors |
| Shuffle slide order | +7.25 | +0.11 | -0.01 |
| Inject one false number | +0.50 | +1.51 | +0.20 |
| Shrink figures by 50% | +0.25 | -0.17 | +0.51 |
| Drop Slides | +2.50 | +0.01 | -0.01 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Sub-stage | Model | Max rounds / calls |
| Planner | Candidate plan generation | GPT-5.5 | 60 |
| Candidate critic and selector | GLM-5.2 | 1 | |
| Plan grounding verifier | GPT-5.5 | 1 | |
| Slide Generator | Slide generation | MiMo-v2.5-Pro | 50 |
| Visual Generator | Visual instruction / prompt generation | MiMo-v2.5-Pro | 15 |
| Image generation | GPT-Image-2 | 1 per generated visual |
| ConfArena score | ||
| Configuration | GPT-5.4 | Gemini-3.8- Flash |
| SlideLab (full) (MiMo-V2.5-Pro) | 0.74 | 0.78 |
| Planner | 0.53 | 0.62 |
| LayoutDebugger | 0.60 | 0.69 |
| Compositor | 0.70 | 0.74 |
| Custom visuals (paper figures only) | 0.69 | 0.70 |
| Per slide | Whole talk | ||||||
| System | Grounding errors | Figure errors | Design | Figure use | Narrative errors | Coverage | Env. score |
| PPTAgent (GPT-5.5) | 1.462 | 0.443 | 3.441 | 3.090 | 5.13 | 0.64 | 0.38 |
| DeepPresenter (GPT-5.5) | 1.882 | 0.316 | 3.725 | 3.368 | 4.77 | 0.753 | 0.30 |
| Kimi Slides | 1.162 | 0.156 | 4.430 | 2.830 | 3.94 | 0.924 | 0.61 |
| Manus | 1.650 | – | 3.990 | – | 3.20 | 0.691 | 0.31 |
| SlideLab (ours) (MiMo-V2.5-Pro, all text stages) | 1.244 | 0.140 | 4.673 | 4.525 | 2.400 | 0.927 | 0.784 |
| Attendee | Focus and Question Trigger |
| Expert reviewer | Focuses on scientific rigor and correctness. Asks a question when a claim, result, or numerical value appears unsupported, inconsistent, or overstated. |
| Learner | Focuses on whether the core ideas, methodology, and motivation are easy to follow. Asks a question when an important concept or step is insufficiently explained. |
| Cross-field attendee | Focuses on accessibility without deep subfield knowledge. Asks a question when jargon or field-specific assumptions are introduced without adequate explanation. |
| Metric | Pearson correlation ( ) |
| Grounding errors | 0.939 |
| Figure errors | 0.952 |
| Design | 0.981 |
| Figure use | 0.917 |
| Coverage | 0.637 |
| Narrative errors | 0.552 |
| Narrative errors | ||
| Evaluator | SlideLab | Kimi Slides |
| Static examiner | 0.90 | 0.80 |
| ConfArena | 2.40 | 3.94 |
| Evaluator condition | Narrative errors per deck | |
| Full ConfArena | 3.88 | |
| Neutral attendee roles | 3.40 | |
| Without Q&A | 2.30 | |
| Examiner only | 1.75 | |
| Stage | Time (s) | Cost ($) | Tokens (M) |
| Planner | 96 | 0.10 | 0.09 |
| Slide Generator | 264 | 0.03 | 0.23 |
| ImageVisualGenerator + Compositor | 261 | 0.42 | 0.32 |
| LayoutDebugger | 122 | 0.03 | 0.05 |
| SlideLab total | 743 | 0.58 | 0.69 |
| DeepPresenter | 1620 | 2.1 | 2.9 |
| Dimension | What annotators are asked |
| Content Grounding | Do the slides match what the paper says? |
| Content Coverage | Does the deck cover the paper’s key contributions? |
| Narrative Structure | Do the slides tell a coherent story? |
| Visual Design | Does it look like a professional conference talk? |
| Information Density | Are slides concise, not walls of text? |
| Figure Usage | Are figures well-chosen and sized correctly? |
| Perturbation | Benchmark | Metric (direction) | Baseline | Perturbed | Detected? | |
| Falsified number | PPTEval | content (1–5, ) | 4.30 | 4.27 | ✓ | |
| SlidesGen-Bench | quiz accuracy (0–1, ) | 0.98 | 0.98 | |||
| PresentBench | material-dep. pass % ( ) | 27.5 | 45.0 | |||
| ConfArena | faithfulness (0–1, ) | 0.58 | 0.57 | ✓ | ||
| Degraded figure | PPTEval | design (1–5, ) | 4.05 | 3.85 | ✓ | |
| SlidesGen-Bench | visual weighted total (0–10, ) | 8.45 | 8.22 | ✓ |