Not just slides - delivery-ready talks
See the codeDeepSlide is not a tool that simply “makes a PPT for you”, but a delivery-first, human-in-the-loop system for complete presentation preparation and delivery.
It helps users go beyond slide generation to support the full workflow of presentation delivery, including narrative planning, logical chain editing, slide generation, script drafting, interactive enhancement, and rehearsal feedback.
🌐 Chinese Version: README_zh.md
We ship an AgentSkills-compatible OpenClaw skill at skills/deepslide-openclaw/:
If your OpenClaw workspace is not the repo root, add skills via skills.load.extraDirs (in ~/.openclaw/openclaw.json)
(ClawHub Link).
A high-quality talk is not mainly determined by whether the static slides “look nice”. What matters more is whether information is organized and delivered under audience cognition and attention constraints—including narrative coherence, timing and pacing control, attention guidance, and rehearsal readiness. In other words, artifact quality ≠ delivery quality.
Figure 1. Comparison with existing methods: DeepSlide targets end-to-end presentation delivery rather than deck authoring only.
To this end, DeepSlide proposes a four-stage end-to-end pipeline: requirement clarification & narrative proposals → logical-chain editing & evidence-grounded generation → interactive enhancement & attention control → rehearsal & evaluation. This shifts presentation preparation from improving artifact quality to improving delivery quality.
recipe/content.tex (slides) and recipe/speech.txt (script)
Figure 2. Overview of the four-stage framework: a closed-loop delivery workflow from requirement clarification to generation/enhancement, and finally rehearsal/evaluation.
To better evaluate from both artifact quality and delivery quality, we developed an LLM-based dual-scoreboard evaluation, and compared against a set of existing methods:
Figure 3. Dual-scoreboard results across 20 domains (Artifact vs. Delivery).
Figure 4. Dual-scoreboard results under mixed role settings (Artifact vs. Delivery).
Figure 5. System UI and key capabilities: logical-chain editing, evidence-grounded generation, interactive enhancement, and a rehearsal loop.
The paper argues that existing “slide agents / generators” typically reduce the cost of deck authoring, but still fail to cover the full burden of talk preparation. There are three main gaps:
DeepSlide’s methodology is: the presenter only needs to lock in high-level decisions (audience, total duration, goals, style intent, narrative skeleton, and emphasis allocation). The system then executes the rest under controllable constraints, forming an iterative delivery loop. This is implemented as four stages:
The core implementation lives in deepslide/, and runtime consists of three services:
deepslide/backend: FastAPI (parsing, generation, compilation, export, evaluation entrypoints)deepslide/frontend: Vite + React (interactive editing, preview, dialog entrypoints)next-ai-draw-io: Next.js (diagrams and draw.io capabilities)DeepSlide/
├── deepslide/
│ ├── backend/
│ ├── frontend/
│ ├── env.md # Model/Agent env vars overview
│ ├── install.sh # Dependency install script (shortcut)
│ ├── start.sh # One-click start for all 3 services
│ ├── stop.sh # One-click stop
│ └── clear.sh # Clear caches/artifacts (dangerous: deletes projects)
├── experiments/ # Evaluation reproduction (dual-scoreboard / ablations)
├── DeepSlide-Arxiv/ # Paper artifact directory (figures/tables/latex)
├── assets/ # README assets (synced from the paper)
└── README_zh.md
xelatex + beamer packages). Strongly recommended to use the provided container/dockerfile so you don’t need to install TeX locally.This repo provides container/dockerfile with TeXLive and Python. You can use Docker to get a ready-to-use TeX compile environment without installing TeX locally:
docker build -t deepslide:latest -f container/dockerfile .
docker run -it --rm \
-v "$(pwd)":/app \
-p 5173:5173 -p 8001:8001 -p 6002:6002 \
deepslide:latest bash
Inside the container, run the same steps below under /app (or directly run deepslide/start.sh).
cd next-ai-draw-io
npm install
cd ..
cd deepslide/backend
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt
cd ../..
cd deepslide/frontend
npm install
cd ../..
You may also use the shortcut script (still recommended to prepare venv/permissions first): bash deepslide/install.sh.
Edit deepslide/.env. See deepslide/env.md for detailed model environment variables.
cd deepslide
bash start.sh
Default endpoints (ports can be changed in .env):
http://127.0.0.1:5173http://127.0.0.1:8001/api/v1http://127.0.0.1:8001/docshttp://127.0.0.1:6002Stop all services:
cd deepslide
bash stop.sh
Replace the key with your own value; do not commit real keys.
# Default text LLM
DEFAULT_MODEL_PLATFORM_TYPE=openai
DEFAULT_MODEL_TYPE=gpt-4o-mini
DEFAULT_MODEL_API_URL=https://api.openai.com/v1
DEFAULT_MODEL_API_KEY=YOUR_API_KEY
# Dev Ports
BACKEND_PORT=8001
FRONTEND_PORT=5173
NEXT_AI_DRAWIO_PORT=6002
DeepSlide supports configuring different provider/model/base_url/api_key for different Agents, so you can use cheaper models for simple steps and stronger models for hard steps. See deepslide/env.md for the full list of fields and Agent names.
deepslide/start.shdeepslide/.envnext-ai-draw-io, backend uvicorn, frontend vitedeepslide/.pids/ for stop/cleanupdeepslide/stop.shdeepslide/clear.sh (Dangerous)Resets runtime state by cleaning caches and generated artifacts (including projects, uploads, and ASR/TTS intermediates). Do not run it if you want to keep project results.
deepslide/install.shInstalls backend/frontend/next-ai-draw-io dependencies (shortcut script).
recipe/content.tex and recipe/speech.txtThe evaluation code is under experiments/. The core idea is a dual-scoreboard: distinguishing static artifact quality (Artifact) vs. delivery quality (Delivery).
It is recommended to create a dedicated venv for evaluation:
python3 -m venv experiments/.venv
source experiments/.venv/bin/activate
pip install --upgrade pip
pip install -r experiments/main/requirements.txt
Copy and fill:
experiments/main/.env.template → experiments/main/.envIt includes:
source experiments/.venv/bin/activate
python experiments/main/run_oneclick.py
Outputs are typically under experiments/main/outputs/ (scores / reports, etc.).
source experiments/.venv/bin/activate
python experiments/role/run_oneclick.py
The ablation entrypoint is experiments/xr/run_eval.py:
source experiments/.venv/bin/activate
python experiments/xr/run_eval.py scan
python experiments/xr/run_eval.py evaluate --judge llm --llm-mode packed
python experiments/xr/run_eval.py report
Notes:
--require-ocr 0 or set EVAL_OCR_MODE=off (metrics depending on OCR will be affected)--require-judge 0 (metrics requiring judge will be skipped)deepslide/start.sh starts this service. Before the first run, make sure you run npm install in next-ai-draw-io/.
The backend TTS logic calls index-tts/index-tts-main (and relies on the uv command). If you need voice preview, follow index-tts/index-tts-main/README.md to install and prepare checkpoints, and ensure uv is available in PATH.
BACKEND_PORT and whether the backend is running; check if PID files exist under deepslide/.pids/xelatex + beamer dependencies are present (or use the Docker environment); check missing fontsEVAL_MODEL_* and DEFAULT_VLM_* in experiments/main/.env.template, or skip via --require-ocr 0/--require-judge 0
|
|
@misc{yang2026deepslideartifactspresentationdelivery,
title={DeepSlide: From Artifacts to Presentation Delivery},
author={Ming Yang and Zhiwei Zhang and Jiahang Li and Haoseng Liu and Yuzheng Cai and Weiguo Zheng},
year={2026},
eprint={2605.15202},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.15202},
}
Python
45.9%
TypeScript
42.3%
JavaScript
5.2%
TeX
5.1%
Not just slides - delivery-ready talks
See the codeDeepSlide is not a tool that simply “makes a PPT for you”, but a delivery-first, human-in-the-loop system for complete presentation preparation and delivery.
It helps users go beyond slide generation to support the full workflow of presentation delivery, including narrative planning, logical chain editing, slide generation, script drafting, interactive enhancement, and rehearsal feedback.
🌐 Chinese Version: README_zh.md
We ship an AgentSkills-compatible OpenClaw skill at skills/deepslide-openclaw/:
If your OpenClaw workspace is not the repo root, add skills via skills.load.extraDirs (in ~/.openclaw/openclaw.json)
(ClawHub Link).
A high-quality talk is not mainly determined by whether the static slides “look nice”. What matters more is whether information is organized and delivered under audience cognition and attention constraints—including narrative coherence, timing and pacing control, attention guidance, and rehearsal readiness. In other words, artifact quality ≠ delivery quality.
Figure 1. Comparison with existing methods: DeepSlide targets end-to-end presentation delivery rather than deck authoring only.
To this end, DeepSlide proposes a four-stage end-to-end pipeline: requirement clarification & narrative proposals → logical-chain editing & evidence-grounded generation → interactive enhancement & attention control → rehearsal & evaluation. This shifts presentation preparation from improving artifact quality to improving delivery quality.
recipe/content.tex (slides) and recipe/speech.txt (script)
Figure 2. Overview of the four-stage framework: a closed-loop delivery workflow from requirement clarification to generation/enhancement, and finally rehearsal/evaluation.
To better evaluate from both artifact quality and delivery quality, we developed an LLM-based dual-scoreboard evaluation, and compared against a set of existing methods:
Figure 3. Dual-scoreboard results across 20 domains (Artifact vs. Delivery).
Figure 4. Dual-scoreboard results under mixed role settings (Artifact vs. Delivery).
Figure 5. System UI and key capabilities: logical-chain editing, evidence-grounded generation, interactive enhancement, and a rehearsal loop.
The paper argues that existing “slide agents / generators” typically reduce the cost of deck authoring, but still fail to cover the full burden of talk preparation. There are three main gaps:
DeepSlide’s methodology is: the presenter only needs to lock in high-level decisions (audience, total duration, goals, style intent, narrative skeleton, and emphasis allocation). The system then executes the rest under controllable constraints, forming an iterative delivery loop. This is implemented as four stages:
The core implementation lives in deepslide/, and runtime consists of three services:
deepslide/backend: FastAPI (parsing, generation, compilation, export, evaluation entrypoints)deepslide/frontend: Vite + React (interactive editing, preview, dialog entrypoints)next-ai-draw-io: Next.js (diagrams and draw.io capabilities)DeepSlide/
├── deepslide/
│ ├── backend/
│ ├── frontend/
│ ├── env.md # Model/Agent env vars overview
│ ├── install.sh # Dependency install script (shortcut)
│ ├── start.sh # One-click start for all 3 services
│ ├── stop.sh # One-click stop
│ └── clear.sh # Clear caches/artifacts (dangerous: deletes projects)
├── experiments/ # Evaluation reproduction (dual-scoreboard / ablations)
├── DeepSlide-Arxiv/ # Paper artifact directory (figures/tables/latex)
├── assets/ # README assets (synced from the paper)
└── README_zh.md
xelatex + beamer packages). Strongly recommended to use the provided container/dockerfile so you don’t need to install TeX locally.This repo provides container/dockerfile with TeXLive and Python. You can use Docker to get a ready-to-use TeX compile environment without installing TeX locally:
docker build -t deepslide:latest -f container/dockerfile .
docker run -it --rm \
-v "$(pwd)":/app \
-p 5173:5173 -p 8001:8001 -p 6002:6002 \
deepslide:latest bash
Inside the container, run the same steps below under /app (or directly run deepslide/start.sh).
cd next-ai-draw-io
npm install
cd ..
cd deepslide/backend
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt
cd ../..
cd deepslide/frontend
npm install
cd ../..
You may also use the shortcut script (still recommended to prepare venv/permissions first): bash deepslide/install.sh.
Edit deepslide/.env. See deepslide/env.md for detailed model environment variables.
cd deepslide
bash start.sh
Default endpoints (ports can be changed in .env):
http://127.0.0.1:5173http://127.0.0.1:8001/api/v1http://127.0.0.1:8001/docshttp://127.0.0.1:6002Stop all services:
cd deepslide
bash stop.sh
Replace the key with your own value; do not commit real keys.
# Default text LLM
DEFAULT_MODEL_PLATFORM_TYPE=openai
DEFAULT_MODEL_TYPE=gpt-4o-mini
DEFAULT_MODEL_API_URL=https://api.openai.com/v1
DEFAULT_MODEL_API_KEY=YOUR_API_KEY
# Dev Ports
BACKEND_PORT=8001
FRONTEND_PORT=5173
NEXT_AI_DRAWIO_PORT=6002
DeepSlide supports configuring different provider/model/base_url/api_key for different Agents, so you can use cheaper models for simple steps and stronger models for hard steps. See deepslide/env.md for the full list of fields and Agent names.
deepslide/start.shdeepslide/.envnext-ai-draw-io, backend uvicorn, frontend vitedeepslide/.pids/ for stop/cleanupdeepslide/stop.shdeepslide/clear.sh (Dangerous)Resets runtime state by cleaning caches and generated artifacts (including projects, uploads, and ASR/TTS intermediates). Do not run it if you want to keep project results.
deepslide/install.shInstalls backend/frontend/next-ai-draw-io dependencies (shortcut script).
recipe/content.tex and recipe/speech.txtThe evaluation code is under experiments/. The core idea is a dual-scoreboard: distinguishing static artifact quality (Artifact) vs. delivery quality (Delivery).
It is recommended to create a dedicated venv for evaluation:
python3 -m venv experiments/.venv
source experiments/.venv/bin/activate
pip install --upgrade pip
pip install -r experiments/main/requirements.txt
Copy and fill:
experiments/main/.env.template → experiments/main/.envIt includes:
source experiments/.venv/bin/activate
python experiments/main/run_oneclick.py
Outputs are typically under experiments/main/outputs/ (scores / reports, etc.).
source experiments/.venv/bin/activate
python experiments/role/run_oneclick.py
The ablation entrypoint is experiments/xr/run_eval.py:
source experiments/.venv/bin/activate
python experiments/xr/run_eval.py scan
python experiments/xr/run_eval.py evaluate --judge llm --llm-mode packed
python experiments/xr/run_eval.py report
Notes:
--require-ocr 0 or set EVAL_OCR_MODE=off (metrics depending on OCR will be affected)--require-judge 0 (metrics requiring judge will be skipped)deepslide/start.sh starts this service. Before the first run, make sure you run npm install in next-ai-draw-io/.
The backend TTS logic calls index-tts/index-tts-main (and relies on the uv command). If you need voice preview, follow index-tts/index-tts-main/README.md to install and prepare checkpoints, and ensure uv is available in PATH.
BACKEND_PORT and whether the backend is running; check if PID files exist under deepslide/.pids/xelatex + beamer dependencies are present (or use the Docker environment); check missing fontsEVAL_MODEL_* and DEFAULT_VLM_* in experiments/main/.env.template, or skip via --require-ocr 0/--require-judge 0
|
|
@misc{yang2026deepslideartifactspresentationdelivery,
title={DeepSlide: From Artifacts to Presentation Delivery},
author={Ming Yang and Zhiwei Zhang and Jiahang Li and Haoseng Liu and Yuzheng Cai and Weiguo Zheng},
year={2026},
eprint={2605.15202},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.15202},
}
Python
45.9%
TypeScript
42.3%
JavaScript
5.2%
TeX
5.1%