A Python toolkit for building vision pipelines: label datasets with VLMs, evaluate annotations with a VLM-as-judge, and train object detection models — all from a few function calls.
This workflow uses separate models for separate roles. Keep them on separate processes/endpoints — do not collapse them into one model.
| Role | Size | What it does | Where it runs |
|---|---|---|---|
| Orchestrator | ~10-12B (thinking) | Drives the workflow, babysits the run, fixes errors, decides next steps. Never labels or judges itself. | The agent/subagent |
| Labeller | ~7B-8B VLM | Generates detections on the dataset (one model, the larger of the workers). | Remote (HF Inference Providers) or local server, it costs $0.5/1k images |
| Judges | 2-5B VLMs (several) | Score/verify the labeller's detections. Use multiple small judges for an ensemble/voting signal. | Remote or local servers |
Why the separation matters:
model_id (and optional base_url), you can run several small judges and
combine their verdicts, rather than relying on a single large model.Concretely, in the label_dataset / judge_labels / train calls:
label_dataset(..., model_id=<7B VLM>) at your labeller endpoint.judge_labels(..., model_id=<4-5B VLM>) at one or more judge endpoints.model_id.This is a uv project:
uv sync # core: label / judge / merge (no torch)
uv sync --extra train # adds torch, transformers, timm, albumentations, … for training
uv sync --extra viz # adds supervision + Roboflow trackers (bbox + tracking viz)
uv sync --extra all # train + viz + dev tooling (pytest, ruff)
Then run anything with uv run (e.g. uv run python -m workflows.vlm_label …).
A requirements.txt is also kept for pip install -r requirements.txt.
Use any VLM to auto-detect objects in images. Works with local image directories or Hugging Face datasets.
from workflows import label_dataset
# Local images → COCO JSON
label_dataset(
source="data/images",
classes=["person", "car", "traffic light"],
output="annotations.json",
)
# HF dataset → HF dataset with detections column
label_dataset(
source="username/my-image-dataset",
classes=["person", "car", "traffic light"],
output="username/my-dataset-labeled",
split="train",
max_samples=1000,
push_to_hub=True,
backend="openai",
base_url="http://localhost:8083/v1", # llama-server
api_key="sk-local",
)
classes is free-form — name whatever the task needs (document regions,
products, road signs, defects, …). Worked use cases live in
examples/, one subfolder each (goal, classes, model-per-role,
commands, outputs); the seed examples/docvqa-media
runs the full label → judge → train pipeline on HF Jobs.
Score each detection and keep only the ones above a threshold.
from workflows import judge_labels
judge_labels(
source="username/my-dataset-labeled",
output="username/my-dataset-judged",
threshold=0.5,
push_to_hub=True,
backend="openai",
base_url="http://localhost:8084/v1", # small 4B judge server
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf",
api_key="sk-local-judge",
)
Fine-tune a detection model on the curated dataset. This is a generalized
version of the HF object-detection tutorial:
lazy preprocessing, optional Albumentations
augmentation, and COCO-style mAP / mAR evaluation via torchmetrics.
from workflows import train
train(
source="username/my-dataset-judged", # or a local COCO directory
model_id="Roboflow/rf-detr-large",
epochs=20,
batch_size=8,
augment=True, # use-case dependent — see note below; off via augment=False
val_split="test", # held-out split for mAP/mAR; None to skip eval
push_to_hub=True,
hub_model_id="username/my-detector",
report_to="trackio", # live metric tracking
)
Supported input formats (auto-detected):
objects column — the standard HF detection layout.detections column — produced by label_dataset /
judge_labels (Pascal-VOC boxes, converted automatically).train/images/ + train/labels.json
(and optional val/).Any AutoModelForObjectDetection checkpoint works (RF-DETR, DETR, etc.); the
default is Roboflow/rf-detr-large. RF-DETR requires transformers>=5.10 and
timm.
Augmentation is use-case dependent. It defaults on (augment=True), but for
clean/uniform imagery or already-tight labels it can hurt — e.g. road-signs
trained markedly better with --no-augment. Assess per dataset; when unsure,
train both ways and compare mAP.
The same functions are exposed as a discoverable, JSON-schema'd tool registry for an in-process orchestrating agent — no MCP server, no per-tool boilerplate. The agent enumerates tools, reads their schemas, and dispatches by name:
import vision_agent as va # or: from tools import get_tools, as_json_schema, call, configure, ToolConfig
# 1. Configure worker endpoints once, by role. Credentials live here — never
# in a tool's schema, so the agent is never asked to fill an API key.
va.configure(
labeller=va.ToolConfig(base_url="https://router.huggingface.co/v1",
model_id="Qwen/Qwen3-VL-8B-Instruct"),
judge=va.ToolConfig(base_url="http://localhost:8084/v1",
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf"),
)
# 2. Discover. get_tools() is torch-free by default (the openai-backed VLM
# tools, the label/judge workflows, and CPU helpers); pass
# include_train=True to also surface the local-GPU tools + `train`.
specs = va.as_json_schema() # [{name, description, parameters}, ...] — plain JSON Schema
print([t.name for t in va.get_tools()])
# 3. Dispatch by name. Hidden backend/model_id/base_url/api_key are injected
# from the role config; pass them explicitly to override per call.
boxes = va.call("vlm_detect", image="photo.jpg", classes=["person", "car"])
configure() accepts default (VLM tools), labeller (label_dataset), and
judge (judge_labels) roles. Any ToolConfig field left None falls through
to the function's own default (model_id), the HF Inference Providers URL
(base_url), or the HF_TOKEN / OPENAI_API_KEY env fallback (api_key). The
same three can be set via VISION_AGENT_BACKEND / VISION_AGENT_MODEL /
VISION_AGENT_BASE_URL.
All VLM-powered tools (vlm_detect, ocr_judge) and
workflows (label_dataset, judge_labels) support two backends:
backend="openai" (recommended for serving)Uses the OpenAI Python client. A single code path that works with:
| Provider | base_url | api_key |
|---|---|---|
| HF Inference Providers | https://router.huggingface.co/v1 (default) | Your HF token |
| vLLM | http://localhost:8000/v1 | Server API key |
| llama-server | http://localhost:8084/v1 | Server API key |
| Any OpenAI-compatible endpoint | Custom URL | Custom key |
from tools import vlm_detect
vlm_detect(
"photo.jpg",
classes=["person", "car"],
backend="openai",
base_url="https://router.huggingface.co/v1",
api_key="hf_...",
model_id="Qwen/Qwen3-VL-8B-Instruct",
)
backend="transformers" (local GPU)Loads the model directly with transformers + torch. Models are cached
after the first load.
from tools import vlm_detect
vlm_detect(
"photo.jpg",
classes=["person", "car"],
backend="transformers",
model_id="Qwen/Qwen2.5-VL-7B-Instruct",
)
| Tool | Description |
|---|---|
detect | RF-DETR closed-set detection (COCO 80 classes) |
instance_segment | RF-DETR-Seg instance segmentation |
segment_from_bbox | SAM3 — bbox prompt to high-quality masks |
segment_from_text | Falcon-Perception — text prompt to zero-shot masks |
estimate_depth | Depth Anything V2 — monocular relative depth |
estimate_pose | Sapiens2 — dense 308-keypoint human pose |
grounded_detect | MM-Grounding-DINO — open-vocabulary detection |
fast_segment | EdgeTAM — lightweight bbox to mask |
ocr | PaddleOCR-VL — vision-language OCR |
vlm_detect | VLM instruction-prompted detection (any VLM) |
ocr_judge | Pairwise OCR quality evaluation with ELO rating |
convert_bbox | Convert bboxes between 6 formats |
validate_annotations | Validate detection annotations for issues |
compute_stats | Statistics for COCO annotation files |
annotate | supervision-backed box + mask visualization (per-class / per-track colours) |
track_video | Roboflow trackers + supervision multi-object tracking (boxes or instance masks) |
| Workflow | Description |
|---|---|
label_dataset | Auto-label images with a VLM for object detection |
judge_labels | Score and filter labels with a VLM-as-judge |
train | Fine-tune RF-DETR on labeled data |
All workflows support both local directories (COCO format) and Hugging Face datasets as input and output.
viz extra)supervision and the
trackers package back the bbox
visualization and multi-object tracking tools. Install with uv sync --extra viz.
Both bridge to the repo's plain detection-dicts
({"label", "score", "box"/"bbox", ...}) via tools.sv_convert, so any
detector here (detect, grounded_detect, vlm_detect) feeds straight in.
from tools import annotate, track_video
# Still image: draw boxes + labels with per-class colours (drop-in for the
# pure-PIL tools.bbox_viz.draw_detections; also takes judge `verdicts`).
annotate("frame.jpg", detections, verdicts=verdicts).save("annotated.png")
# Video: detect (or instance-segment) every frame, associate with a Roboflow
# tracker, and write an annotated copy with stable per-object colours, ids,
# motion trails — and masks when a segmentation detector is used.
track_video("clip.mp4", detector="rfdetr", tracker="bytetrack") # boxes
track_video("clip.mp4", detector="rfdetr-seg") # tracked instance masks
track_video("clip.mp4", detector="falcon", classes=["red car"]) # open-vocab masks
Per-frame detectors: rfdetr (RF-DETR boxes), grounded (MM-Grounding-DINO,
open-vocab boxes), vlm (VLM-prompted boxes), rfdetr-seg (RF-DETR-Seg instance
masks) and falcon (Falcon-Perception, open-vocab masks). Trackers associate on
boxes (derived from masks for the segmentation detectors) and the masks ride
along, so tracked instances keep a stable colour + id frame to frame.
# Annotate one image from a detections JSON file
python -m tools.sv_viz frame.jpg -d dets.json --out annotated.png
# Track objects through a video (bytetrack | sort | ocsort | botsort)
python -m tools.track_video clip.mp4 --detector grounded --classes "car,person" \
--tracker bytetrack --out tracked.mp4
track_video needs both the viz extra and a detector from the train extra
(torch); the still-image annotate runs on a CPU-only viz install.
Each workflow can also be run from the command line:
# Label (7-8B labeller, here via HF Inference Providers)
python -m workflows.vlm_label \
--source username/my-image-dataset \
--classes "person,car,traffic light" \
--output username/my-dataset-labeled --push-to-hub \
--backend openai --base-url https://router.huggingface.co/v1 \
--api-key hf_... --model Qwen/Qwen3-VL-8B-Instruct \
--split train --max-samples 100
# Judge (small 4B judge, here on a local llama-server)
python -m workflows.vlm_judge \
--source username/my-dataset-labeled \
--output username/my-dataset-judged --push-to-hub \
--backend openai --base-url http://localhost:8084/v1 \
--api-key sk-local-judge --model Qwen3-VL-4B-Instruct-Q8_0.gguf \
--threshold 0.5
# Train (mAP/mAR eval on a held-out split, push to the Hub)
python -m workflows.train_rfdetr \
--source username/my-dataset-judged --val-split test \
--model Roboflow/rf-detr-large --epochs 20 --batch-size 8 \
--output-dir checkpoints/my-detector
Each role uses its own model (see "Recommended architecture" above). The
labeller is the larger worker; the judge is a small, cheap verifier. The
~12B orchestrator drives this script and is never used as a model_id
here. (For a fully worked multi-model run on HF Jobs, see
jobs/README.md.)
Serve a small judge locally with llama-server (runs alongside the orchestrator with little VRAM):
llama-server \
--model Qwen3-VL-4B-Instruct-Q8_0.gguf \
--mmproj mmproj-F16.gguf \
--host 0.0.0.0 --port 8084 \
--n-gpu-layers 99 --ctx-size 8192 \
--api-key sk-local-judge
Then run the pipeline from Python:
from workflows import label_dataset, judge_labels, train
# Labeller: ~7-8B VLM via HF Inference Providers (remote)
LABELLER = dict(
backend="openai",
base_url="https://router.huggingface.co/v1",
api_key="hf_...",
model_id="Qwen/Qwen3-VL-8B-Instruct",
)
# Judge: small ~4B VLM on a local llama-server
JUDGE = dict(
backend="openai",
base_url="http://localhost:8084/v1",
api_key="sk-local-judge",
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf",
)
# Step 1: label
label_dataset(
source="username/my-image-dataset",
classes=["person", "car", "traffic light"],
output="username/my-dataset-labeled",
split="train",
max_samples=1000,
push_to_hub=True,
**LABELLER,
)
# Step 2: judge
judge_labels(
source="username/my-dataset-labeled",
output="username/my-dataset-judged",
threshold=0.5,
push_to_hub=True,
**JUDGE,
)
# Step 3: train
train(
source="username/my-dataset-judged",
epochs=20,
batch_size=4,
)
MIT
14 commits
1 commits
Python
100.0%
A Python toolkit for building vision pipelines: label datasets with VLMs, evaluate annotations with a VLM-as-judge, and train object detection models — all from a few function calls.
This workflow uses separate models for separate roles. Keep them on separate processes/endpoints — do not collapse them into one model.
| Role | Size | What it does | Where it runs |
|---|---|---|---|
| Orchestrator | ~10-12B (thinking) | Drives the workflow, babysits the run, fixes errors, decides next steps. Never labels or judges itself. | The agent/subagent |
| Labeller | ~7B-8B VLM | Generates detections on the dataset (one model, the larger of the workers). | Remote (HF Inference Providers) or local server, it costs $0.5/1k images |
| Judges | 2-5B VLMs (several) | Score/verify the labeller's detections. Use multiple small judges for an ensemble/voting signal. | Remote or local servers |
Why the separation matters:
model_id (and optional base_url), you can run several small judges and
combine their verdicts, rather than relying on a single large model.Concretely, in the label_dataset / judge_labels / train calls:
label_dataset(..., model_id=<7B VLM>) at your labeller endpoint.judge_labels(..., model_id=<4-5B VLM>) at one or more judge endpoints.model_id.This is a uv project:
uv sync # core: label / judge / merge (no torch)
uv sync --extra train # adds torch, transformers, timm, albumentations, … for training
uv sync --extra viz # adds supervision + Roboflow trackers (bbox + tracking viz)
uv sync --extra all # train + viz + dev tooling (pytest, ruff)
Then run anything with uv run (e.g. uv run python -m workflows.vlm_label …).
A requirements.txt is also kept for pip install -r requirements.txt.
Use any VLM to auto-detect objects in images. Works with local image directories or Hugging Face datasets.
from workflows import label_dataset
# Local images → COCO JSON
label_dataset(
source="data/images",
classes=["person", "car", "traffic light"],
output="annotations.json",
)
# HF dataset → HF dataset with detections column
label_dataset(
source="username/my-image-dataset",
classes=["person", "car", "traffic light"],
output="username/my-dataset-labeled",
split="train",
max_samples=1000,
push_to_hub=True,
backend="openai",
base_url="http://localhost:8083/v1", # llama-server
api_key="sk-local",
)
classes is free-form — name whatever the task needs (document regions,
products, road signs, defects, …). Worked use cases live in
examples/, one subfolder each (goal, classes, model-per-role,
commands, outputs); the seed examples/docvqa-media
runs the full label → judge → train pipeline on HF Jobs.
Score each detection and keep only the ones above a threshold.
from workflows import judge_labels
judge_labels(
source="username/my-dataset-labeled",
output="username/my-dataset-judged",
threshold=0.5,
push_to_hub=True,
backend="openai",
base_url="http://localhost:8084/v1", # small 4B judge server
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf",
api_key="sk-local-judge",
)
Fine-tune a detection model on the curated dataset. This is a generalized
version of the HF object-detection tutorial:
lazy preprocessing, optional Albumentations
augmentation, and COCO-style mAP / mAR evaluation via torchmetrics.
from workflows import train
train(
source="username/my-dataset-judged", # or a local COCO directory
model_id="Roboflow/rf-detr-large",
epochs=20,
batch_size=8,
augment=True, # use-case dependent — see note below; off via augment=False
val_split="test", # held-out split for mAP/mAR; None to skip eval
push_to_hub=True,
hub_model_id="username/my-detector",
report_to="trackio", # live metric tracking
)
Supported input formats (auto-detected):
objects column — the standard HF detection layout.detections column — produced by label_dataset /
judge_labels (Pascal-VOC boxes, converted automatically).train/images/ + train/labels.json
(and optional val/).Any AutoModelForObjectDetection checkpoint works (RF-DETR, DETR, etc.); the
default is Roboflow/rf-detr-large. RF-DETR requires transformers>=5.10 and
timm.
Augmentation is use-case dependent. It defaults on (augment=True), but for
clean/uniform imagery or already-tight labels it can hurt — e.g. road-signs
trained markedly better with --no-augment. Assess per dataset; when unsure,
train both ways and compare mAP.
The same functions are exposed as a discoverable, JSON-schema'd tool registry for an in-process orchestrating agent — no MCP server, no per-tool boilerplate. The agent enumerates tools, reads their schemas, and dispatches by name:
import vision_agent as va # or: from tools import get_tools, as_json_schema, call, configure, ToolConfig
# 1. Configure worker endpoints once, by role. Credentials live here — never
# in a tool's schema, so the agent is never asked to fill an API key.
va.configure(
labeller=va.ToolConfig(base_url="https://router.huggingface.co/v1",
model_id="Qwen/Qwen3-VL-8B-Instruct"),
judge=va.ToolConfig(base_url="http://localhost:8084/v1",
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf"),
)
# 2. Discover. get_tools() is torch-free by default (the openai-backed VLM
# tools, the label/judge workflows, and CPU helpers); pass
# include_train=True to also surface the local-GPU tools + `train`.
specs = va.as_json_schema() # [{name, description, parameters}, ...] — plain JSON Schema
print([t.name for t in va.get_tools()])
# 3. Dispatch by name. Hidden backend/model_id/base_url/api_key are injected
# from the role config; pass them explicitly to override per call.
boxes = va.call("vlm_detect", image="photo.jpg", classes=["person", "car"])
configure() accepts default (VLM tools), labeller (label_dataset), and
judge (judge_labels) roles. Any ToolConfig field left None falls through
to the function's own default (model_id), the HF Inference Providers URL
(base_url), or the HF_TOKEN / OPENAI_API_KEY env fallback (api_key). The
same three can be set via VISION_AGENT_BACKEND / VISION_AGENT_MODEL /
VISION_AGENT_BASE_URL.
All VLM-powered tools (vlm_detect, ocr_judge) and
workflows (label_dataset, judge_labels) support two backends:
backend="openai" (recommended for serving)Uses the OpenAI Python client. A single code path that works with:
| Provider | base_url | api_key |
|---|---|---|
| HF Inference Providers | https://router.huggingface.co/v1 (default) | Your HF token |
| vLLM | http://localhost:8000/v1 | Server API key |
| llama-server | http://localhost:8084/v1 | Server API key |
| Any OpenAI-compatible endpoint | Custom URL | Custom key |
from tools import vlm_detect
vlm_detect(
"photo.jpg",
classes=["person", "car"],
backend="openai",
base_url="https://router.huggingface.co/v1",
api_key="hf_...",
model_id="Qwen/Qwen3-VL-8B-Instruct",
)
backend="transformers" (local GPU)Loads the model directly with transformers + torch. Models are cached
after the first load.
from tools import vlm_detect
vlm_detect(
"photo.jpg",
classes=["person", "car"],
backend="transformers",
model_id="Qwen/Qwen2.5-VL-7B-Instruct",
)
| Tool | Description |
|---|---|
detect | RF-DETR closed-set detection (COCO 80 classes) |
instance_segment | RF-DETR-Seg instance segmentation |
segment_from_bbox | SAM3 — bbox prompt to high-quality masks |
segment_from_text | Falcon-Perception — text prompt to zero-shot masks |
estimate_depth | Depth Anything V2 — monocular relative depth |
estimate_pose | Sapiens2 — dense 308-keypoint human pose |
grounded_detect | MM-Grounding-DINO — open-vocabulary detection |
fast_segment | EdgeTAM — lightweight bbox to mask |
ocr | PaddleOCR-VL — vision-language OCR |
vlm_detect | VLM instruction-prompted detection (any VLM) |
ocr_judge | Pairwise OCR quality evaluation with ELO rating |
convert_bbox | Convert bboxes between 6 formats |
validate_annotations | Validate detection annotations for issues |
compute_stats | Statistics for COCO annotation files |
annotate | supervision-backed box + mask visualization (per-class / per-track colours) |
track_video | Roboflow trackers + supervision multi-object tracking (boxes or instance masks) |
| Workflow | Description |
|---|---|
label_dataset | Auto-label images with a VLM for object detection |
judge_labels | Score and filter labels with a VLM-as-judge |
train | Fine-tune RF-DETR on labeled data |
All workflows support both local directories (COCO format) and Hugging Face datasets as input and output.
viz extra)supervision and the
trackers package back the bbox
visualization and multi-object tracking tools. Install with uv sync --extra viz.
Both bridge to the repo's plain detection-dicts
({"label", "score", "box"/"bbox", ...}) via tools.sv_convert, so any
detector here (detect, grounded_detect, vlm_detect) feeds straight in.
from tools import annotate, track_video
# Still image: draw boxes + labels with per-class colours (drop-in for the
# pure-PIL tools.bbox_viz.draw_detections; also takes judge `verdicts`).
annotate("frame.jpg", detections, verdicts=verdicts).save("annotated.png")
# Video: detect (or instance-segment) every frame, associate with a Roboflow
# tracker, and write an annotated copy with stable per-object colours, ids,
# motion trails — and masks when a segmentation detector is used.
track_video("clip.mp4", detector="rfdetr", tracker="bytetrack") # boxes
track_video("clip.mp4", detector="rfdetr-seg") # tracked instance masks
track_video("clip.mp4", detector="falcon", classes=["red car"]) # open-vocab masks
Per-frame detectors: rfdetr (RF-DETR boxes), grounded (MM-Grounding-DINO,
open-vocab boxes), vlm (VLM-prompted boxes), rfdetr-seg (RF-DETR-Seg instance
masks) and falcon (Falcon-Perception, open-vocab masks). Trackers associate on
boxes (derived from masks for the segmentation detectors) and the masks ride
along, so tracked instances keep a stable colour + id frame to frame.
# Annotate one image from a detections JSON file
python -m tools.sv_viz frame.jpg -d dets.json --out annotated.png
# Track objects through a video (bytetrack | sort | ocsort | botsort)
python -m tools.track_video clip.mp4 --detector grounded --classes "car,person" \
--tracker bytetrack --out tracked.mp4
track_video needs both the viz extra and a detector from the train extra
(torch); the still-image annotate runs on a CPU-only viz install.
Each workflow can also be run from the command line:
# Label (7-8B labeller, here via HF Inference Providers)
python -m workflows.vlm_label \
--source username/my-image-dataset \
--classes "person,car,traffic light" \
--output username/my-dataset-labeled --push-to-hub \
--backend openai --base-url https://router.huggingface.co/v1 \
--api-key hf_... --model Qwen/Qwen3-VL-8B-Instruct \
--split train --max-samples 100
# Judge (small 4B judge, here on a local llama-server)
python -m workflows.vlm_judge \
--source username/my-dataset-labeled \
--output username/my-dataset-judged --push-to-hub \
--backend openai --base-url http://localhost:8084/v1 \
--api-key sk-local-judge --model Qwen3-VL-4B-Instruct-Q8_0.gguf \
--threshold 0.5
# Train (mAP/mAR eval on a held-out split, push to the Hub)
python -m workflows.train_rfdetr \
--source username/my-dataset-judged --val-split test \
--model Roboflow/rf-detr-large --epochs 20 --batch-size 8 \
--output-dir checkpoints/my-detector
Each role uses its own model (see "Recommended architecture" above). The
labeller is the larger worker; the judge is a small, cheap verifier. The
~12B orchestrator drives this script and is never used as a model_id
here. (For a fully worked multi-model run on HF Jobs, see
jobs/README.md.)
Serve a small judge locally with llama-server (runs alongside the orchestrator with little VRAM):
llama-server \
--model Qwen3-VL-4B-Instruct-Q8_0.gguf \
--mmproj mmproj-F16.gguf \
--host 0.0.0.0 --port 8084 \
--n-gpu-layers 99 --ctx-size 8192 \
--api-key sk-local-judge
Then run the pipeline from Python:
from workflows import label_dataset, judge_labels, train
# Labeller: ~7-8B VLM via HF Inference Providers (remote)
LABELLER = dict(
backend="openai",
base_url="https://router.huggingface.co/v1",
api_key="hf_...",
model_id="Qwen/Qwen3-VL-8B-Instruct",
)
# Judge: small ~4B VLM on a local llama-server
JUDGE = dict(
backend="openai",
base_url="http://localhost:8084/v1",
api_key="sk-local-judge",
model_id="Qwen3-VL-4B-Instruct-Q8_0.gguf",
)
# Step 1: label
label_dataset(
source="username/my-image-dataset",
classes=["person", "car", "traffic light"],
output="username/my-dataset-labeled",
split="train",
max_samples=1000,
push_to_hub=True,
**LABELLER,
)
# Step 2: judge
judge_labels(
source="username/my-dataset-labeled",
output="username/my-dataset-judged",
threshold=0.5,
push_to_hub=True,
**JUDGE,
)
# Step 3: train
train(
source="username/my-dataset-judged",
epochs=20,
batch_size=4,
)
MIT
14 commits
1 commits
Python
100.0%