CVHub520/X-AnyLabeling

X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

Python

10,546

1,487 commits

updated Sep 19, 2026

See the code

README

X-AnyLabeling point cloud annotation

X-AnyLabeling interface

🥳 What's New

  • 2026-09-18: Add 3D point cloud annotation, with per-point semantic and instance labeling, camera image reference, and calibrated point overlays.
  • 2026-08-19: Add support for image tagging, with tag creation, editing, reordering, and batch deletion.
  • 2026-08-12: Add support for D-FINE-seg instance segmentation models.
  • 2026-08-08: Add support for the RT-DETRv2-OBB rotated object detection model.
  • 2026-08-08: Add the Magic Wand tool for quickly creating polygons from contiguous color regions.
  • 2026-08-05: Release X-AnyLabeling v4.0.0.
  • For more details, please refer to the CHANGELOG

Introduction

X-AnyLabeling is a lightweight, efficient, and unified cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It combines versatile built-in tools, automated labeling workflows, state-of-the-art deep learning models, and flexible multi-format import and export. For remote inference, X-AnyLabeling-Server provides a lightweight, extensible backend for connecting custom models and compute resources.

Key Features

  • Unified support for annotating and processing text, image, video, point cloud, and multimodal data.
  • Covers tasks such as image classification, object detection, instance segmentation, pose estimation, oriented object detection, multi-object tracking, optical character recognition, lane annotation, image captioning, visual question answering, document parsing, and 3D point clouds.
  • Provides polygons, rectangles, cuboids, rotated boxes, quadrilaterals, circles, lines, polylines, points, masks, and task-specific tools for text detection, text recognition, and KIE.
  • Integrates a wide range of state-of-the-art deep learning models for AI-assisted annotation, automated labeling, and batch dataset prediction.
  • Supports both local and remote inference through engines and serving frameworks such as ONNX Runtime, TensorRT, OpenCV DNN, vLLM, and SGLang.
  • Supports importing and exporting formats such as COCO, VOC, YOLO, DOTA, MOT, MASK, PPOCR, MMGD, VLM-R1, and ShareGPT.
  • Runs on Windows, Linux, and macOS, with interfaces available in English, Simplified Chinese, Japanese, and Korean.
  • Supports custom model integration, flexible extension, and secondary development.

Model library

Task CategorySupported Models
🖼️ Image ClassificationYOLOv5-Cls, YOLOv8-Cls, YOLO11-Cls, InternImage, PULC
🎯 Object DetectionYOLOv5/6/7/8/9/10, YOLO11/12/26, YOLOX, YOLO-NAS, D-FINE, DAMO-YOLO, Gold_YOLO, RT-DETR, RF-DETR, DEIMv2
🖌️ Instance SegmentationYOLOv5-Seg, YOLOv8-Seg, YOLO11-Seg, YOLO26-Seg, Hyper-YOLO-Seg, RF-DETR-Seg, D-FINE-seg
🏃 Pose EstimationYOLOv8-Pose, YOLO11-Pose, YOLO26-Pose, DWPose, RTMO
😀 Face EstimationSCRFD, YOLOv6Lite-Face
👣 TrackingTrackTrack, Bot-SORT, ByteTrack, SAM2/3-Video
🔄 Rotated Object DetectionYOLOv5-Obb, YOLOv8-Obb, YOLO11-Obb, YOLO26-Obb, RT-DETRv2-OBB
📏 Depth EstimationDepth Anything
🧩 Segment AnythingSAM 1/2/3, SAM-HQ, SAM-Med2D, EdgeSAM, EfficientViT-SAM, MobileSAM
✂️ Image MattingRMBG 1.4/2.0
💡 ProposalUPN
🏷️ TaggingRAM, RAM++
📄 OCRPP-OCRv4, PP-OCRv5, PP-OCRv6
🧾 Layout AnalysisPP-DocLayoutV3
📑 Document ParsingPaddleOCR-VL, PaddleOCR-VL-1.6
🗣️ Vision Foundation ModelsRex-Omni, Florence2
👁️ Vision Language ModelsQwen3-VL, Gemini, ChatGPT, GLM
🛣️ Lane DetectionCLRNet
🔢 Object CountingCountGD, GeCO, GeCo2
📍 GroundingGrounding DINO, YOLO-World, YOLOE, SAM 3, LocateAnything
📚 Other👉 model_zoo 👈

Docs

  1. Remote Inference Service
  2. Installation & Quickstart
  3. Usage
  4. Command Line Interface
  5. Customize a model
  6. Chatbot
  7. VQA
  8. Image Classifier
  9. Video Classifier
  10. Document Parsing and Intelligent Text Recognition
  11. 3D Point Cloud Annotation

Examples

Contribute

We believe in open collaboration! X‑AnyLabeling continues to grow with the support of the community. Whether you're fixing bugs, improving documentation, or adding new features, your contributions make a real impact.

To get started, please read our Contributing Guide and make sure to agree to the Contributor License Agreement (CLA) before submitting a pull request.

If you find this project helpful, please consider giving it a ⭐️ star! Have questions or suggestions? Open an issue or email us at cv_hub@163.com.

A huge thank you 🙏 to everyone helping to make X‑AnyLabeling better.

License

This project is licensed under the GNU General Public License v3.0. You may use, modify, and redistribute the software, including for commercial purposes, provided that you comply with the terms of the license.

X-AnyLabeling is an actively maintained open-source project. Your sponsorship helps support feature development, model integration, documentation, and community support.

Sponsor the X-AnyLabeling project

Click the image above to visit the sponsorship page.

Acknowledgement

I extend my heartfelt thanks to the developers and contributors of AnyLabeling, LabelMe, LabelImg, roLabelImg, PPOCRLabel and CVAT, whose work has been crucial to the success of this project.

Citing

If you use this software in your research, please cite it as below:

@misc{X-AnyLabeling,
  year = {2023},
  author = {Wei Wang},
  publisher = {Github},
  organization = {CVHub},
  journal = {Github repository},
  title = {X-AnyLabeling: A Unified Desktop Platform for AI-Assisted Data Annotation},
  howpublished = {\url{https://github.com/CVHub520/X-AnyLabeling}}
}
artificial-intelligence
clip
computer-vision
deep-learning
groundingdino
image-annotation-tool
image-classification
image-labeling-tool
image-matting
instance-segmentation
machine-learning
object-detection
ocr
onnxruntime
paddlepaddle
pose-estimation
rotated-object-detection
sam
vision-language-model
yolo

Contributors

(top 30 of 50)

CVHub520

1,359 commits

zhixuwei

17 commits

chevydream

12 commits

PairZhu

10 commits

CVHub520/X-AnyLabeling

X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

Python

10,546

1,487 commits

updated Sep 19, 2026

See the code

README

X-AnyLabeling point cloud annotation

X-AnyLabeling interface

🥳 What's New

  • 2026-09-18: Add 3D point cloud annotation, with per-point semantic and instance labeling, camera image reference, and calibrated point overlays.
  • 2026-08-19: Add support for image tagging, with tag creation, editing, reordering, and batch deletion.
  • 2026-08-12: Add support for D-FINE-seg instance segmentation models.
  • 2026-08-08: Add support for the RT-DETRv2-OBB rotated object detection model.
  • 2026-08-08: Add the Magic Wand tool for quickly creating polygons from contiguous color regions.
  • 2026-08-05: Release X-AnyLabeling v4.0.0.
  • For more details, please refer to the CHANGELOG

Introduction

X-AnyLabeling is a lightweight, efficient, and unified cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It combines versatile built-in tools, automated labeling workflows, state-of-the-art deep learning models, and flexible multi-format import and export. For remote inference, X-AnyLabeling-Server provides a lightweight, extensible backend for connecting custom models and compute resources.

Key Features

  • Unified support for annotating and processing text, image, video, point cloud, and multimodal data.
  • Covers tasks such as image classification, object detection, instance segmentation, pose estimation, oriented object detection, multi-object tracking, optical character recognition, lane annotation, image captioning, visual question answering, document parsing, and 3D point clouds.
  • Provides polygons, rectangles, cuboids, rotated boxes, quadrilaterals, circles, lines, polylines, points, masks, and task-specific tools for text detection, text recognition, and KIE.
  • Integrates a wide range of state-of-the-art deep learning models for AI-assisted annotation, automated labeling, and batch dataset prediction.
  • Supports both local and remote inference through engines and serving frameworks such as ONNX Runtime, TensorRT, OpenCV DNN, vLLM, and SGLang.
  • Supports importing and exporting formats such as COCO, VOC, YOLO, DOTA, MOT, MASK, PPOCR, MMGD, VLM-R1, and ShareGPT.
  • Runs on Windows, Linux, and macOS, with interfaces available in English, Simplified Chinese, Japanese, and Korean.
  • Supports custom model integration, flexible extension, and secondary development.

Model library

Task CategorySupported Models
🖼️ Image ClassificationYOLOv5-Cls, YOLOv8-Cls, YOLO11-Cls, InternImage, PULC
🎯 Object DetectionYOLOv5/6/7/8/9/10, YOLO11/12/26, YOLOX, YOLO-NAS, D-FINE, DAMO-YOLO, Gold_YOLO, RT-DETR, RF-DETR, DEIMv2
🖌️ Instance SegmentationYOLOv5-Seg, YOLOv8-Seg, YOLO11-Seg, YOLO26-Seg, Hyper-YOLO-Seg, RF-DETR-Seg, D-FINE-seg
🏃 Pose EstimationYOLOv8-Pose, YOLO11-Pose, YOLO26-Pose, DWPose, RTMO
😀 Face EstimationSCRFD, YOLOv6Lite-Face
👣 TrackingTrackTrack, Bot-SORT, ByteTrack, SAM2/3-Video
🔄 Rotated Object DetectionYOLOv5-Obb, YOLOv8-Obb, YOLO11-Obb, YOLO26-Obb, RT-DETRv2-OBB
📏 Depth EstimationDepth Anything
🧩 Segment AnythingSAM 1/2/3, SAM-HQ, SAM-Med2D, EdgeSAM, EfficientViT-SAM, MobileSAM
✂️ Image MattingRMBG 1.4/2.0
💡 ProposalUPN
🏷️ TaggingRAM, RAM++
📄 OCRPP-OCRv4, PP-OCRv5, PP-OCRv6
🧾 Layout AnalysisPP-DocLayoutV3
📑 Document ParsingPaddleOCR-VL, PaddleOCR-VL-1.6
🗣️ Vision Foundation ModelsRex-Omni, Florence2
👁️ Vision Language ModelsQwen3-VL, Gemini, ChatGPT, GLM
🛣️ Lane DetectionCLRNet
🔢 Object CountingCountGD, GeCO, GeCo2
📍 GroundingGrounding DINO, YOLO-World, YOLOE, SAM 3, LocateAnything
📚 Other👉 model_zoo 👈

Docs

  1. Remote Inference Service
  2. Installation & Quickstart
  3. Usage
  4. Command Line Interface
  5. Customize a model
  6. Chatbot
  7. VQA
  8. Image Classifier
  9. Video Classifier
  10. Document Parsing and Intelligent Text Recognition
  11. 3D Point Cloud Annotation

Examples

Contribute

We believe in open collaboration! X‑AnyLabeling continues to grow with the support of the community. Whether you're fixing bugs, improving documentation, or adding new features, your contributions make a real impact.

To get started, please read our Contributing Guide and make sure to agree to the Contributor License Agreement (CLA) before submitting a pull request.

If you find this project helpful, please consider giving it a ⭐️ star! Have questions or suggestions? Open an issue or email us at cv_hub@163.com.

A huge thank you 🙏 to everyone helping to make X‑AnyLabeling better.

License

This project is licensed under the GNU General Public License v3.0. You may use, modify, and redistribute the software, including for commercial purposes, provided that you comply with the terms of the license.

X-AnyLabeling is an actively maintained open-source project. Your sponsorship helps support feature development, model integration, documentation, and community support.

Sponsor the X-AnyLabeling project

Click the image above to visit the sponsorship page.

Acknowledgement

I extend my heartfelt thanks to the developers and contributors of AnyLabeling, LabelMe, LabelImg, roLabelImg, PPOCRLabel and CVAT, whose work has been crucial to the success of this project.

Citing

If you use this software in your research, please cite it as below:

@misc{X-AnyLabeling,
  year = {2023},
  author = {Wei Wang},
  publisher = {Github},
  organization = {CVHub},
  journal = {Github repository},
  title = {X-AnyLabeling: A Unified Desktop Platform for AI-Assisted Data Annotation},
  howpublished = {\url{https://github.com/CVHub520/X-AnyLabeling}}
}
artificial-intelligence
clip
computer-vision
deep-learning
groundingdino
image-annotation-tool
image-classification
image-labeling-tool
image-matting
instance-segmentation
machine-learning
object-detection
ocr
onnxruntime
paddlepaddle
pose-estimation
rotated-object-detection
sam
vision-language-model
yolo

Contributors

(top 30 of 50)

CVHub520

1,359 commits

zhixuwei

17 commits

chevydream

12 commits

PairZhu

10 commits

Languages

Python

98.8%

Cuda

1.1%