An offline AI-powered video analysis tool with object detection (YOLO), image captioning (BLIP), speech transcription (Whisper), audio event detection (PANNs), and AI-generated summaries (LLMs via Ollama). It ensures privacy and offline use with a user-friendly GUI.
Python
105
12 commits
updated Jun 14, 2026
Offline, privacy-first AI video analysis. Runs entirely on your local machine — no cloud, no data upload, no telemetry.
Analyzes a video file through a local AI pipeline and produces structured outputs:
| Stage | Technology | Default |
|---|---|---|
| Object detection | VisionServeX D-FINE | Required — primary backend |
| Frame sampling | Built-in (adaptive/scene-change) | Enabled |
| Scene captioning | BLIP (Salesforce) | Optional — needs [full] |
| Speech transcription | Whisper (OpenAI) | Optional — needs [full] |
| Audio events | PANNs CNN14 | Optional — needs [full] |
| LLM summarization | Ollama (any local model) | Optional — needs Ollama |
All processing is local. Nothing is sent to any server.
# 1. Clone and install (detection only)
git clone https://github.com/arashsajjadi/ai-powered-video-analyzer.git
cd ai-powered-video-analyzer
python -m pip install -U pip
python -m pip install -e ".[vision]"
# 2. Check your environment
ai-video-analyzer doctor
# 3. Analyze a video
ai-video-analyzer analyze "/path/to/video.mp4" --preset balanced
python -m pip install -e ".[vision]"
Installs: frame sampler, VisionServeX D-FINE backend, CLI. No YOLO, no cloud dependencies.
python -m pip install -r pip_requirements.txt
Adds: Whisper, BLIP, PANNs, PyTorch, librosa, Ollama client.
python -m pip install -e ".[dev]"
pytest -q
# Ubuntu / Debian
sudo apt install ffmpeg
# macOS
brew install ffmpeg
# Install
curl -fsSL https://ollama.com/install.sh | sh
# Start and pull a model
ollama serve
ollama pull phi4:latest
# Check that everything is installed
ai-video-analyzer doctor
# Run detection only (fast, no optional deps needed)
ai-video-analyzer analyze "/path/to/video.mp4" \
--preset balanced \
--no-captioning \
--no-audio-events \
--no-summarization
# Full analysis (requires full install + Ollama)
ai-video-analyzer analyze "/path/to/video.mp4" \
--preset balanced \
--whisper-model base \
--ollama-model phi4:latest \
--output-dir ./results
ai-video-analyzer doctor
Checks: Python version, OpenCV, VisionServeX, D-FINE model registry, ffmpeg, PyTorch/GPU, Whisper, BLIP, PANNs, Ollama, moviepy, Tesseract.
Required dependencies exit with ✗. Optional dependencies show ✓ with an install hint when missing.
Exits 0 if all required dependencies are present.
VisionServeX D-FINE models (COCO-80 classes, benchmarked on RTX 5080):
| Preset | Model | ~ms/frame | Notes |
|---|---|---|---|
fast | dfine-n | 17–21 | Speed-first; good accuracy |
balanced | dfine-s | 17–27 | Default — best all-round |
quality | dfine-m | 21–32 | Highest COCO accuracy |
quality+ | dfine-l | higher | Maximum accuracy, slowest |
ai-video-analyzer analyze video.mp4 --preset fast
ai-video-analyzer analyze video.mp4 --preset quality
# Override with a specific model ID
ai-video-analyzer analyze video.mp4 --model dfine-s
# List all available models
ai-video-analyzer list-models
Known limitation: COCO-80 does not include fire, smoke, or weather classes. D-FINE will return approximate visual matches (e.g. food textures for fire/smoke frames). The BLIP captioning stage handles these cases in natural language.
Each run writes to --output-dir (default: current directory):
| File | Contents |
|---|---|
<video>_analysis.json | Full structured data: detections, captions, transcript, timings, preset |
<video>_analysis.md | Human-readable Markdown report with tables |
report.txt | Plain-text legacy report |
The JSON report includes: preset, frame_strategy, timings, top labels, and limitations.
# Single preset
ai-video-analyzer benchmark "/path/to/video.mp4" --preset balanced
# Compare fast / balanced / quality side-by-side
ai-video-analyzer benchmark "/path/to/video.mp4" --compare
Benchmark output:
Video duration : 8.2s
Frames selected : 8 (sampling: 0.08s)
Model load+warmup : 0.52s
Detection runtime : 0.15s (18.6ms/frame, 53.8fps)
Total detections : 48
Top labels : dog(9), person(7)
Model : dfine-s
Real-world results in reports/benchmarks/.
When Ollama is running, the pipeline generates a factual evidence-grounded summary.
# Start Ollama
ollama serve
ollama pull phi4:latest
# Run with summarization
ai-video-analyzer analyze video.mp4 --preset balanced --ollama-model phi4:latest
# List installed models
ai-video-analyzer analyze video.mp4 --list-ollama-models
The summary prompt separates observed facts from plausible interpretation, and says "Insufficient evidence" instead of guessing when data is weak.
pip install 'visionservex[hf,rfdetr]'
# or
pip install -e ".[vision]"
sudo apt install ffmpeg # Ubuntu/Debian
brew install ffmpeg # macOS
python -c "import torch; print(torch.cuda.is_available())"
ai-video-analyzer analyze video.mp4 --device cpu
Run doctor to check the detection backend:
ai-video-analyzer doctor
If dfine-s shows as unavailable, try listing what is registered:
ai-video-analyzer list-models
ai-video-analyzer doctor
The original video_processing.py interface is preserved:
python video_processing.py --video "/path/to/video.mp4" --save
python video_processing.py --video video.mp4 --preset fast --no-summarization
The Tkinter GUI is also preserved:
python video_processing_gui.py
ai-video-analyzer gui
Legacy optional backend: YOLO-based detection is available as --backend legacy_yolo
with the [legacy-yolo] extra (pip install -e ".[legacy-yolo]").
This is not recommended for new use. VisionServeX D-FINE is the default and supported backend.
Video file
│
├─ Adaptive frame sampling (scene + motion aware)
│
├─ Object detection → VisionServeX D-FINE (per frame)
├─ Scene captioning → BLIP (optional)
│
├─ Audio extraction
│ ├─ Transcription → Whisper (optional)
│ └─ Audio events → PANNs CNN14 (optional)
│
└─ LLM summarization → Ollama (optional)
│
└─ AnalysisReport
├─ <video>_analysis.json
├─ <video>_analysis.md
└─ report.txt
git clone https://github.com/arashsajjadi/ai-powered-video-analyzer.git
cd ai-powered-video-analyzer
python -m pip install -e ".[dev]"
pytest -q
ruff check .
See docs/LLM_AGENT_GUIDE.md for the complete agent guide.
Quick reference:
python -m pip install -e ".[vision]"pytest -qai-video-analyzer doctor/home/arash/PycharmProjects/VisionServeX is read-onlyai_powered_video_analyzer/config.py — AnalysisConfig dataclassai_powered_video_analyzer/core.py — analyze_video(config)ai_powered_video_analyzer/backends/visionservex_backend.pyMIT — see LICENSE.
Developed by Arash Sajjadi, University of Saskatchewan.
Python
100.0%
An offline AI-powered video analysis tool with object detection (YOLO), image captioning (BLIP), speech transcription (Whisper), audio event detection (PANNs), and AI-generated summaries (LLMs via Ollama). It ensures privacy and offline use with a user-friendly GUI.
Python
105
12 commits
updated Jun 14, 2026
Offline, privacy-first AI video analysis. Runs entirely on your local machine — no cloud, no data upload, no telemetry.
Analyzes a video file through a local AI pipeline and produces structured outputs:
| Stage | Technology | Default |
|---|---|---|
| Object detection | VisionServeX D-FINE | Required — primary backend |
| Frame sampling | Built-in (adaptive/scene-change) | Enabled |
| Scene captioning | BLIP (Salesforce) | Optional — needs [full] |
| Speech transcription | Whisper (OpenAI) | Optional — needs [full] |
| Audio events | PANNs CNN14 | Optional — needs [full] |
| LLM summarization | Ollama (any local model) | Optional — needs Ollama |
All processing is local. Nothing is sent to any server.
# 1. Clone and install (detection only)
git clone https://github.com/arashsajjadi/ai-powered-video-analyzer.git
cd ai-powered-video-analyzer
python -m pip install -U pip
python -m pip install -e ".[vision]"
# 2. Check your environment
ai-video-analyzer doctor
# 3. Analyze a video
ai-video-analyzer analyze "/path/to/video.mp4" --preset balanced
python -m pip install -e ".[vision]"
Installs: frame sampler, VisionServeX D-FINE backend, CLI. No YOLO, no cloud dependencies.
python -m pip install -r pip_requirements.txt
Adds: Whisper, BLIP, PANNs, PyTorch, librosa, Ollama client.
python -m pip install -e ".[dev]"
pytest -q
# Ubuntu / Debian
sudo apt install ffmpeg
# macOS
brew install ffmpeg
# Install
curl -fsSL https://ollama.com/install.sh | sh
# Start and pull a model
ollama serve
ollama pull phi4:latest
# Check that everything is installed
ai-video-analyzer doctor
# Run detection only (fast, no optional deps needed)
ai-video-analyzer analyze "/path/to/video.mp4" \
--preset balanced \
--no-captioning \
--no-audio-events \
--no-summarization
# Full analysis (requires full install + Ollama)
ai-video-analyzer analyze "/path/to/video.mp4" \
--preset balanced \
--whisper-model base \
--ollama-model phi4:latest \
--output-dir ./results
ai-video-analyzer doctor
Checks: Python version, OpenCV, VisionServeX, D-FINE model registry, ffmpeg, PyTorch/GPU, Whisper, BLIP, PANNs, Ollama, moviepy, Tesseract.
Required dependencies exit with ✗. Optional dependencies show ✓ with an install hint when missing.
Exits 0 if all required dependencies are present.
VisionServeX D-FINE models (COCO-80 classes, benchmarked on RTX 5080):
| Preset | Model | ~ms/frame | Notes |
|---|---|---|---|
fast | dfine-n | 17–21 | Speed-first; good accuracy |
balanced | dfine-s | 17–27 | Default — best all-round |
quality | dfine-m | 21–32 | Highest COCO accuracy |
quality+ | dfine-l | higher | Maximum accuracy, slowest |
ai-video-analyzer analyze video.mp4 --preset fast
ai-video-analyzer analyze video.mp4 --preset quality
# Override with a specific model ID
ai-video-analyzer analyze video.mp4 --model dfine-s
# List all available models
ai-video-analyzer list-models
Known limitation: COCO-80 does not include fire, smoke, or weather classes. D-FINE will return approximate visual matches (e.g. food textures for fire/smoke frames). The BLIP captioning stage handles these cases in natural language.
Each run writes to --output-dir (default: current directory):
| File | Contents |
|---|---|
<video>_analysis.json | Full structured data: detections, captions, transcript, timings, preset |
<video>_analysis.md | Human-readable Markdown report with tables |
report.txt | Plain-text legacy report |
The JSON report includes: preset, frame_strategy, timings, top labels, and limitations.
# Single preset
ai-video-analyzer benchmark "/path/to/video.mp4" --preset balanced
# Compare fast / balanced / quality side-by-side
ai-video-analyzer benchmark "/path/to/video.mp4" --compare
Benchmark output:
Video duration : 8.2s
Frames selected : 8 (sampling: 0.08s)
Model load+warmup : 0.52s
Detection runtime : 0.15s (18.6ms/frame, 53.8fps)
Total detections : 48
Top labels : dog(9), person(7)
Model : dfine-s
Real-world results in reports/benchmarks/.
When Ollama is running, the pipeline generates a factual evidence-grounded summary.
# Start Ollama
ollama serve
ollama pull phi4:latest
# Run with summarization
ai-video-analyzer analyze video.mp4 --preset balanced --ollama-model phi4:latest
# List installed models
ai-video-analyzer analyze video.mp4 --list-ollama-models
The summary prompt separates observed facts from plausible interpretation, and says "Insufficient evidence" instead of guessing when data is weak.
pip install 'visionservex[hf,rfdetr]'
# or
pip install -e ".[vision]"
sudo apt install ffmpeg # Ubuntu/Debian
brew install ffmpeg # macOS
python -c "import torch; print(torch.cuda.is_available())"
ai-video-analyzer analyze video.mp4 --device cpu
Run doctor to check the detection backend:
ai-video-analyzer doctor
If dfine-s shows as unavailable, try listing what is registered:
ai-video-analyzer list-models
ai-video-analyzer doctor
The original video_processing.py interface is preserved:
python video_processing.py --video "/path/to/video.mp4" --save
python video_processing.py --video video.mp4 --preset fast --no-summarization
The Tkinter GUI is also preserved:
python video_processing_gui.py
ai-video-analyzer gui
Legacy optional backend: YOLO-based detection is available as --backend legacy_yolo
with the [legacy-yolo] extra (pip install -e ".[legacy-yolo]").
This is not recommended for new use. VisionServeX D-FINE is the default and supported backend.
Video file
│
├─ Adaptive frame sampling (scene + motion aware)
│
├─ Object detection → VisionServeX D-FINE (per frame)
├─ Scene captioning → BLIP (optional)
│
├─ Audio extraction
│ ├─ Transcription → Whisper (optional)
│ └─ Audio events → PANNs CNN14 (optional)
│
└─ LLM summarization → Ollama (optional)
│
└─ AnalysisReport
├─ <video>_analysis.json
├─ <video>_analysis.md
└─ report.txt
git clone https://github.com/arashsajjadi/ai-powered-video-analyzer.git
cd ai-powered-video-analyzer
python -m pip install -e ".[dev]"
pytest -q
ruff check .
See docs/LLM_AGENT_GUIDE.md for the complete agent guide.
Quick reference:
python -m pip install -e ".[vision]"pytest -qai-video-analyzer doctor/home/arash/PycharmProjects/VisionServeX is read-onlyai_powered_video_analyzer/config.py — AnalysisConfig dataclassai_powered_video_analyzer/core.py — analyze_video(config)ai_powered_video_analyzer/backends/visionservex_backend.pyMIT — see LICENSE.
Developed by Arash Sajjadi, University of Saskatchewan.
Python
100.0%