ailia-ai/ailia-models

The collection of pre-trained, state-of-the-art AI models for ailia SDK

2,389

stars

6,608

commits

Python

primary language

Sep 6, 2026

updated

action-recognition
anomaly-detection
audio-processing
background-removal
crowd-counting
deep-learning
embeddings
face-detection
face-recognition
fashion-ai
gan
hand-detection
image-classification
image-segmentation
llm
neural-network
object-detection
object-recognition
object-tracking
pose-estimation

README

GitHub stars PyPI Python Platform

The collection of pre-trained, state-of-the-art AI models: 418 models covering object detection, speech recognition, image generation, LLMs and more — all runnable from the same simple CLI.

Tutorial · チュートリアル · Google Colaboratory · Documentation · deepwiki · Update history

Every model works the same way: no arguments needed, weights download automatically.

pip3 install ailia
git clone https://github.com/ailia-ai/ailia-models
cd ailia-models
pip3 install -r requirements.txt
cd object_detection/yolox
python3 yolox.py

Models

419 models are available. Use 🔍 Search models to find a model by name.

CategoryModel list
Action recognitionva-cnn, st-gcn, mars, ax_action_recognition, driver-action-recognition-adas, action_clip
Anomaly detectionmahalanobisad, spade-pytorch, padim, patchcore, glass
Audio language modelqwen_audio
Audio processingAudio classification: crnn_audio_classification, audioset_tagging_cnn, transformer-cnn-emotion-recognition, microsoft clap, clap
Music enhancement: hifigan, deep music enhancer
Music generation: pytorch_wavenet
Noise reduction: rnnoise, voicefilter, unet_source_separation, demucs, dtln, voicesplit, audiosep
Phoneme alignment: narabas
Pitch detection: crepe
Speaker diarization: pyannote-audio, auto_speech, wespeaker
Speech to text: deepspeech2, whisper, reazon_speech, distil-whisper, sensevoice, reazon_speech2, kotoba-whisper, lite-whisper
Text to speech: pytorch-dc-tts, tacotron2, vall-e-x, Bert-VITS2, gpt-sovits, gpt-sovits-v2, cosyvoice2, gpt-sovits-v3, gpt-sovits-v2-pro, qwen3-tts
Voice activity detection: silero-vad
Voice conversion: rvc
Autonomous drivingbevformer, segformer, uniad
Background removaldeep-image-matting, indexnet, U-2-Net, u2net-portrait-matting, u2net-human-seg, cascade_psp, rembg, gfm, modnet, background_matting_v2, dis_seg
Crowd countingcrowdcount-cascaded-mtl, c-3-framework
Deep fashionfashionai-key-points-detection, person-attributes-recognition-crossroad, clothing-detection, mmfashion, mmfashion_tryon, mmfashion_retrieval
Depth estimationfcrn-depthprediction, monodepth2, fast-depth, midas, hitnet, lap-depth, mobilestereonet, crestereo, zoe_depth, depth_anything, depth_anything_v2, depth_pro, depth_anything_v3
DiffusionText to image: latent-diffusion-txt2img, stable-diffusion-txt2img, anything_v3, control_net, sdxl, latent-consistency-models, sd-turbo, sdxl-turbo, depth_anything_controlnet, latentsync
Text to audio: riffusion
Others: latent-diffusion-inpainting, latent-diffusion-superresolution, DA-CLIP, marigold
Face detectionmtcnn, yolov1-face, face-detection-adas, retinaface, blazeface, yolov3-face, face-mask-detection, dbface, anime-face-detector
Face identificationfacenet_pytorch, insightface, vggface2, arcface, cosface
Face recognitionAge gender estimation: face_classification, age-gender-recognition-retail, mivolo
Emotion recognition: ferplus, hsemotion
Gaze estimation: gazeml, mediapipe_iris, gazelle, ax_gaze_estimation
Head pose estimation: hopenet, 6d_repnet, L2CS_Net, 6d_repnet_360
Keypoint detection: face_alignment, prnet, facemesh, facial_feature, 3ddfa, facemesh_v2
Others: face-anti-spoofing, ax_facial_features
Face restorationgfpgan, codeformer
Face swappingdeepfacelive, sber-swap, facefusion
Feature extractiondinov3
Frame interpolationcain, rife, flavr, film
Generative adversarial networkspytorch-gan, lipgan, council-gan, sam, encoder4editing, restyle-encoder, SadTalker, live_portrait
Hand detectionhand_detection_pytorch, yolov3-hand, blazepalm
Hand recognitionhand3d, v2v-posenet, minimal-hand, blazehand, hands_segmentation_pytorch
Image captioningillustration2vec, image_captioning_pytorch, blip2
Image classificationCNN: alexnet, vgg16, googlenet, resnet18, resnet50, inceptionv3, inceptionv4, wide_resnet50, mobilenetv2, mobilenetv3, efficientnet, efficientnetv2, imagenet21k, mlp_mixer, volo, convnext, mobileone
Transformer: vit, clip, swin-transformer, japanese-clip, japanese-stable-clip-vit-l-16, siglip-multilingual, clip-japanese-base, siglip2
Specific task: weather-prediction-from-image, partialconv
Image inpaintinginpainting-with-partial-conv, deepfillv2, inpainting_gmcnn, 3d-photo-inpainting, lama
Image manipulationcolorization, cnngeometric_pytorch, style2paints, deblur_gan, pytorch-superpoint, noise2noise, dfe, illnet, dewarpnet, deep_white_balance, u2net_portrait, invertible_denoising_network, dfm, fbcnn, dehamer, lightglue, docshadow
Image quality assessmentaesthetic-predictor
Image restorationnafnet
Image segmentationpytorch-fcn, pytorch-enet, tusimple-DUC, pytorch-unet, deeplabv3, pspnet-hair-segmentation, swiftnet, hrnet_segmentation, hair_segmentation, paddleseg, human_part_segmentation, semantic-segmentation-mobilenet-v3, suim, yet-another-anime-segmenter, dense_prediction_transformers, group_vit, pp_liteseg, anime-segmentation, yolov8-seg, segment-anything, grounded_sam, fast_sam, mobile_sam, edge_sam, segment-anything-2, yolov11-seg, segment-anything-3.1
Landmark classificationplaces365, landmarks_classifier_asia
Line segment detectiondexined, mlsd
Low light image enhancementagllnet, drbn_skf
Natural language processingBert: bert, bert_maskedlm, bert_question_answering
Embedding: sentence_transformers_japanese, multilingual-e5, glucose, qwen3-embedding, ruri-v3, embeddinggemma
Error corrector: bert_insert_punctuation, bertjsc, t5_whisper_medical
Grapheme to phoneme: g2p_en, g2pw, soundchoice-g2p
Named entity recognition: bert_ner, t5_base_japanese_ner, bert_ner_japanese
Reranker: cross_encoder_mmarco, japanese-reranker-cross-encoder, ruri-v3-reranker
Sentence generation: gpt2, rinna_gpt2
Sentiment analysis: bert_sentiment_analysis, bert_tweets_sentiment
Summarize: bert_sum_ext, presumm, t5_base_japanese_title_generation, t5_base_summarization
Translation: fugumt-en-ja, fugumt-ja-en
Zero shot classification: bert_zero_shot_classification, multilingual-minilmv2
Network intrusion detectionbert-network-packet-flow-header-payload, falcon-adapter-network-packet
Neural renderingnerf, TripoSR
NSFW detectorclip-based-nsfw-detector
Object detectionCNN: yolov1-tiny, yolov2, yolov2-tiny, maskrcnn, yolov3, yolov3-tiny, mobilenet_ssd, m2det, centernet, yolact, efficientdet, pedestrian_detection, crowd_det, yolov4, yolov4-tiny, yolov5, poly_yolo, nanodet, yolor, yolox, picodet, yolox-ti-lite, yolov7, fastest-det, yolov, yolov6, damo_yolo, yolov8, yolox_body_head_hand_face, yolov9, yolov10, yolov11, yolov12
Transformer: detr, glip, dab-detr, detic, groundingdino, rt-detr-v2
Specific target: traffic-sign-detection, sku110k-densedet, footandball, qrcode_wechatqrcode, mobile_object_localizer, layout_parsing
Object detection 3d3d_bbox, d4lcn, egonet, mediapipe_objectron, 3d-object-detection.pytorch, did_m3d
Object trackingdeepsort, person_reid_baseline_pytorch, abd_net, deepsort_vehicle, qd-3dt, centroids-reid, siam-mot, bytetrack, strong_sort, samurai
Optical flow estimationraft, cotracker3
Point segmentationpointnet_pytorch
Pose estimationopenpose, posenet, pose_resnet, lightweight-human-pose-estimation, animalpose, efficientpose, blazepose, mediapipe_holistic, movenet, ap-10k, e2pose
Pose estimation 3dpose-hg-3d, 3d-pose-baseline, lightweight-human-pose-estimation-3d, 3dmppe_posenet, gast, blazepose-fullbody, mediapipe_pose_world_landmarks
Road detectionroad-segmentation-adas, codes-for-lane-detection, ultra-fast-lane-detection, polylanenet, roneld, lstr, yolop, cdnet, hybridnets
Rotation predictionrotnet
Style transferadain, pix2pixHD, beauty_gan, psgan, animeganv2, EleGANt
Super resolutionsrresnet, edsr, han, real-esrgan, swinir, rcan-it, Hat, SPAN
Text detectioneast, pixel_link, craft_pytorch
Text recognitionetl, crnn.pytorch, deep-text-recognition-benchmark, easyocr, paddleocr, donut, ndlocr_text_recognition, paddleocr_v3
Time-series forecastinginformer2020, timesfm, moirai, chronos2
Vehicle recognitionvehicle-attributes-recognition-barrier, vehicle-license-plate-detection-barrier
Vision language modelllava, florence2, mobilevlm, llava-jp, qwen2_vl, qwen2.5_vl, qwen3_vl
Commercial modelacculus-pose

About ailia SDK

ailia SDK is a cross-platform, high-speed inference SDK for AI. It supports Windows, Mac, Linux, iOS, Android, Jetson, and Raspberry Pi with GPU acceleration via Vulkan and Metal. Bindings are available for C++, Python, Unity (C#), Kotlin, Rust, and Flutter.

Other platforms

Prototype with ailia MODELS (Python), then deploy to production.

Contact

Contributors

(top 30 of 49)

kyakuno

3,760 commits

ooe1123

1,028 commits

sngyo

424 commits

claude

148 commits

ailia-ai/ailia-models

The collection of pre-trained, state-of-the-art AI models for ailia SDK

2,389

stars

6,608

commits

Python

primary language

Sep 6, 2026

updated

action-recognition
anomaly-detection
audio-processing
background-removal
crowd-counting
deep-learning
embeddings
face-detection
face-recognition
fashion-ai
gan
hand-detection
image-classification
image-segmentation
llm
neural-network
object-detection
object-recognition
object-tracking
pose-estimation

README

GitHub stars PyPI Python Platform

The collection of pre-trained, state-of-the-art AI models: 418 models covering object detection, speech recognition, image generation, LLMs and more — all runnable from the same simple CLI.

Tutorial · チュートリアル · Google Colaboratory · Documentation · deepwiki · Update history

Every model works the same way: no arguments needed, weights download automatically.

pip3 install ailia
git clone https://github.com/ailia-ai/ailia-models
cd ailia-models
pip3 install -r requirements.txt
cd object_detection/yolox
python3 yolox.py

Models

419 models are available. Use 🔍 Search models to find a model by name.

CategoryModel list
Action recognitionva-cnn, st-gcn, mars, ax_action_recognition, driver-action-recognition-adas, action_clip
Anomaly detectionmahalanobisad, spade-pytorch, padim, patchcore, glass
Audio language modelqwen_audio
Audio processingAudio classification: crnn_audio_classification, audioset_tagging_cnn, transformer-cnn-emotion-recognition, microsoft clap, clap
Music enhancement: hifigan, deep music enhancer
Music generation: pytorch_wavenet
Noise reduction: rnnoise, voicefilter, unet_source_separation, demucs, dtln, voicesplit, audiosep
Phoneme alignment: narabas
Pitch detection: crepe
Speaker diarization: pyannote-audio, auto_speech, wespeaker
Speech to text: deepspeech2, whisper, reazon_speech, distil-whisper, sensevoice, reazon_speech2, kotoba-whisper, lite-whisper
Text to speech: pytorch-dc-tts, tacotron2, vall-e-x, Bert-VITS2, gpt-sovits, gpt-sovits-v2, cosyvoice2, gpt-sovits-v3, gpt-sovits-v2-pro, qwen3-tts
Voice activity detection: silero-vad
Voice conversion: rvc
Autonomous drivingbevformer, segformer, uniad
Background removaldeep-image-matting, indexnet, U-2-Net, u2net-portrait-matting, u2net-human-seg, cascade_psp, rembg, gfm, modnet, background_matting_v2, dis_seg
Crowd countingcrowdcount-cascaded-mtl, c-3-framework
Deep fashionfashionai-key-points-detection, person-attributes-recognition-crossroad, clothing-detection, mmfashion, mmfashion_tryon, mmfashion_retrieval
Depth estimationfcrn-depthprediction, monodepth2, fast-depth, midas, hitnet, lap-depth, mobilestereonet, crestereo, zoe_depth, depth_anything, depth_anything_v2, depth_pro, depth_anything_v3
DiffusionText to image: latent-diffusion-txt2img, stable-diffusion-txt2img, anything_v3, control_net, sdxl, latent-consistency-models, sd-turbo, sdxl-turbo, depth_anything_controlnet, latentsync
Text to audio: riffusion
Others: latent-diffusion-inpainting, latent-diffusion-superresolution, DA-CLIP, marigold
Face detectionmtcnn, yolov1-face, face-detection-adas, retinaface, blazeface, yolov3-face, face-mask-detection, dbface, anime-face-detector
Face identificationfacenet_pytorch, insightface, vggface2, arcface, cosface
Face recognitionAge gender estimation: face_classification, age-gender-recognition-retail, mivolo
Emotion recognition: ferplus, hsemotion
Gaze estimation: gazeml, mediapipe_iris, gazelle, ax_gaze_estimation
Head pose estimation: hopenet, 6d_repnet, L2CS_Net, 6d_repnet_360
Keypoint detection: face_alignment, prnet, facemesh, facial_feature, 3ddfa, facemesh_v2
Others: face-anti-spoofing, ax_facial_features
Face restorationgfpgan, codeformer
Face swappingdeepfacelive, sber-swap, facefusion
Feature extractiondinov3
Frame interpolationcain, rife, flavr, film
Generative adversarial networkspytorch-gan, lipgan, council-gan, sam, encoder4editing, restyle-encoder, SadTalker, live_portrait
Hand detectionhand_detection_pytorch, yolov3-hand, blazepalm
Hand recognitionhand3d, v2v-posenet, minimal-hand, blazehand, hands_segmentation_pytorch
Image captioningillustration2vec, image_captioning_pytorch, blip2
Image classificationCNN: alexnet, vgg16, googlenet, resnet18, resnet50, inceptionv3, inceptionv4, wide_resnet50, mobilenetv2, mobilenetv3, efficientnet, efficientnetv2, imagenet21k, mlp_mixer, volo, convnext, mobileone
Transformer: vit, clip, swin-transformer, japanese-clip, japanese-stable-clip-vit-l-16, siglip-multilingual, clip-japanese-base, siglip2
Specific task: weather-prediction-from-image, partialconv
Image inpaintinginpainting-with-partial-conv, deepfillv2, inpainting_gmcnn, 3d-photo-inpainting, lama
Image manipulationcolorization, cnngeometric_pytorch, style2paints, deblur_gan, pytorch-superpoint, noise2noise, dfe, illnet, dewarpnet, deep_white_balance, u2net_portrait, invertible_denoising_network, dfm, fbcnn, dehamer, lightglue, docshadow
Image quality assessmentaesthetic-predictor
Image restorationnafnet
Image segmentationpytorch-fcn, pytorch-enet, tusimple-DUC, pytorch-unet, deeplabv3, pspnet-hair-segmentation, swiftnet, hrnet_segmentation, hair_segmentation, paddleseg, human_part_segmentation, semantic-segmentation-mobilenet-v3, suim, yet-another-anime-segmenter, dense_prediction_transformers, group_vit, pp_liteseg, anime-segmentation, yolov8-seg, segment-anything, grounded_sam, fast_sam, mobile_sam, edge_sam, segment-anything-2, yolov11-seg, segment-anything-3.1
Landmark classificationplaces365, landmarks_classifier_asia
Line segment detectiondexined, mlsd
Low light image enhancementagllnet, drbn_skf
Natural language processingBert: bert, bert_maskedlm, bert_question_answering
Embedding: sentence_transformers_japanese, multilingual-e5, glucose, qwen3-embedding, ruri-v3, embeddinggemma
Error corrector: bert_insert_punctuation, bertjsc, t5_whisper_medical
Grapheme to phoneme: g2p_en, g2pw, soundchoice-g2p
Named entity recognition: bert_ner, t5_base_japanese_ner, bert_ner_japanese
Reranker: cross_encoder_mmarco, japanese-reranker-cross-encoder, ruri-v3-reranker
Sentence generation: gpt2, rinna_gpt2
Sentiment analysis: bert_sentiment_analysis, bert_tweets_sentiment
Summarize: bert_sum_ext, presumm, t5_base_japanese_title_generation, t5_base_summarization
Translation: fugumt-en-ja, fugumt-ja-en
Zero shot classification: bert_zero_shot_classification, multilingual-minilmv2
Network intrusion detectionbert-network-packet-flow-header-payload, falcon-adapter-network-packet
Neural renderingnerf, TripoSR
NSFW detectorclip-based-nsfw-detector
Object detectionCNN: yolov1-tiny, yolov2, yolov2-tiny, maskrcnn, yolov3, yolov3-tiny, mobilenet_ssd, m2det, centernet, yolact, efficientdet, pedestrian_detection, crowd_det, yolov4, yolov4-tiny, yolov5, poly_yolo, nanodet, yolor, yolox, picodet, yolox-ti-lite, yolov7, fastest-det, yolov, yolov6, damo_yolo, yolov8, yolox_body_head_hand_face, yolov9, yolov10, yolov11, yolov12
Transformer: detr, glip, dab-detr, detic, groundingdino, rt-detr-v2
Specific target: traffic-sign-detection, sku110k-densedet, footandball, qrcode_wechatqrcode, mobile_object_localizer, layout_parsing
Object detection 3d3d_bbox, d4lcn, egonet, mediapipe_objectron, 3d-object-detection.pytorch, did_m3d
Object trackingdeepsort, person_reid_baseline_pytorch, abd_net, deepsort_vehicle, qd-3dt, centroids-reid, siam-mot, bytetrack, strong_sort, samurai
Optical flow estimationraft, cotracker3
Point segmentationpointnet_pytorch
Pose estimationopenpose, posenet, pose_resnet, lightweight-human-pose-estimation, animalpose, efficientpose, blazepose, mediapipe_holistic, movenet, ap-10k, e2pose
Pose estimation 3dpose-hg-3d, 3d-pose-baseline, lightweight-human-pose-estimation-3d, 3dmppe_posenet, gast, blazepose-fullbody, mediapipe_pose_world_landmarks
Road detectionroad-segmentation-adas, codes-for-lane-detection, ultra-fast-lane-detection, polylanenet, roneld, lstr, yolop, cdnet, hybridnets
Rotation predictionrotnet
Style transferadain, pix2pixHD, beauty_gan, psgan, animeganv2, EleGANt
Super resolutionsrresnet, edsr, han, real-esrgan, swinir, rcan-it, Hat, SPAN
Text detectioneast, pixel_link, craft_pytorch
Text recognitionetl, crnn.pytorch, deep-text-recognition-benchmark, easyocr, paddleocr, donut, ndlocr_text_recognition, paddleocr_v3
Time-series forecastinginformer2020, timesfm, moirai, chronos2
Vehicle recognitionvehicle-attributes-recognition-barrier, vehicle-license-plate-detection-barrier
Vision language modelllava, florence2, mobilevlm, llava-jp, qwen2_vl, qwen2.5_vl, qwen3_vl
Commercial modelacculus-pose

About ailia SDK

ailia SDK is a cross-platform, high-speed inference SDK for AI. It supports Windows, Mac, Linux, iOS, Android, Jetson, and Raspberry Pi with GPU acceleration via Vulkan and Metal. Bindings are available for C++, Python, Unity (C#), Kotlin, Rust, and Flutter.

Other platforms

Prototype with ailia MODELS (Python), then deploy to production.

Contact

Contributors

(top 30 of 49)

kyakuno

3,760 commits

ooe1123

1,028 commits

sngyo

424 commits

claude

148 commits

Languages

Python

95.5%

Jupyter Notebook

4.0%