 | Action recognition | va-cnn, st-gcn, mars, ax_action_recognition, driver-action-recognition-adas, action_clip |
 | Anomaly detection | mahalanobisad, spade-pytorch, padim, patchcore, glass |
| Audio language model | qwen_audio |
| Audio processing | Audio classification: crnn_audio_classification, audioset_tagging_cnn, transformer-cnn-emotion-recognition, microsoft clap, clap Music enhancement: hifigan, deep music enhancer Music generation: pytorch_wavenet Noise reduction: rnnoise, voicefilter, unet_source_separation, demucs, dtln, voicesplit, audiosep Phoneme alignment: narabas Pitch detection: crepe Speaker diarization: pyannote-audio, auto_speech, wespeaker Speech to text: deepspeech2, whisper, reazon_speech, distil-whisper, sensevoice, reazon_speech2, kotoba-whisper, lite-whisper Text to speech: pytorch-dc-tts, tacotron2, vall-e-x, Bert-VITS2, gpt-sovits, gpt-sovits-v2, cosyvoice2, gpt-sovits-v3, gpt-sovits-v2-pro, qwen3-tts Voice activity detection: silero-vad Voice conversion: rvc |
 | Autonomous driving | bevformer, segformer, uniad |
 | Background removal | deep-image-matting, indexnet, U-2-Net, u2net-portrait-matting, u2net-human-seg, cascade_psp, rembg, gfm, modnet, background_matting_v2, dis_seg |
 | Crowd counting | crowdcount-cascaded-mtl, c-3-framework |
 | Deep fashion | fashionai-key-points-detection, person-attributes-recognition-crossroad, clothing-detection, mmfashion, mmfashion_tryon, mmfashion_retrieval |
 | Depth estimation | fcrn-depthprediction, monodepth2, fast-depth, midas, hitnet, lap-depth, mobilestereonet, crestereo, zoe_depth, depth_anything, depth_anything_v2, depth_pro, depth_anything_v3 |
 | Diffusion | Text to image: latent-diffusion-txt2img, stable-diffusion-txt2img, anything_v3, control_net, sdxl, latent-consistency-models, sd-turbo, sdxl-turbo, depth_anything_controlnet, latentsync Text to audio: riffusion Others: latent-diffusion-inpainting, latent-diffusion-superresolution, DA-CLIP, marigold |
 | Face detection | mtcnn, yolov1-face, face-detection-adas, retinaface, blazeface, yolov3-face, face-mask-detection, dbface, anime-face-detector |
 | Face identification | facenet_pytorch, insightface, vggface2, arcface, cosface |
 | Face recognition | Age gender estimation: face_classification, age-gender-recognition-retail, mivolo Emotion recognition: ferplus, hsemotion Gaze estimation: gazeml, mediapipe_iris, gazelle, ax_gaze_estimation Head pose estimation: hopenet, 6d_repnet, L2CS_Net, 6d_repnet_360 Keypoint detection: face_alignment, prnet, facemesh, facial_feature, 3ddfa, facemesh_v2 Others: face-anti-spoofing, ax_facial_features |
 | Face restoration | gfpgan, codeformer |
 | Face swapping | deepfacelive, sber-swap, facefusion |
 | Feature extraction | dinov3 |
 | Frame interpolation | cain, rife, flavr, film |
 | Generative adversarial networks | pytorch-gan, lipgan, council-gan, sam, encoder4editing, restyle-encoder, SadTalker, live_portrait |
 | Hand detection | hand_detection_pytorch, yolov3-hand, blazepalm |
 | Hand recognition | hand3d, v2v-posenet, minimal-hand, blazehand, hands_segmentation_pytorch |
 | Image captioning | illustration2vec, image_captioning_pytorch, blip2 |
 | Image classification | CNN: alexnet, vgg16, googlenet, resnet18, resnet50, inceptionv3, inceptionv4, wide_resnet50, mobilenetv2, mobilenetv3, efficientnet, efficientnetv2, imagenet21k, mlp_mixer, volo, convnext, mobileone Transformer: vit, clip, swin-transformer, japanese-clip, japanese-stable-clip-vit-l-16, siglip-multilingual, clip-japanese-base, siglip2 Specific task: weather-prediction-from-image, partialconv |
 | Image inpainting | inpainting-with-partial-conv, deepfillv2, inpainting_gmcnn, 3d-photo-inpainting, lama |
 | Image manipulation | colorization, cnngeometric_pytorch, style2paints, deblur_gan, pytorch-superpoint, noise2noise, dfe, illnet, dewarpnet, deep_white_balance, u2net_portrait, invertible_denoising_network, dfm, fbcnn, dehamer, lightglue, docshadow |
 | Image quality assessment | aesthetic-predictor |
 | Image restoration | nafnet |
 | Image segmentation | pytorch-fcn, pytorch-enet, tusimple-DUC, pytorch-unet, deeplabv3, pspnet-hair-segmentation, swiftnet, hrnet_segmentation, hair_segmentation, paddleseg, human_part_segmentation, semantic-segmentation-mobilenet-v3, suim, yet-another-anime-segmenter, dense_prediction_transformers, group_vit, pp_liteseg, anime-segmentation, yolov8-seg, segment-anything, grounded_sam, fast_sam, mobile_sam, edge_sam, segment-anything-2, yolov11-seg, segment-anything-3.1 |
 | Landmark classification | places365, landmarks_classifier_asia |
 | Line segment detection | dexined, mlsd |
 | Low light image enhancement | agllnet, drbn_skf |
| Natural language processing | Bert: bert, bert_maskedlm, bert_question_answering Embedding: sentence_transformers_japanese, multilingual-e5, glucose, qwen3-embedding, ruri-v3, embeddinggemma Error corrector: bert_insert_punctuation, bertjsc, t5_whisper_medical Grapheme to phoneme: g2p_en, g2pw, soundchoice-g2p Named entity recognition: bert_ner, t5_base_japanese_ner, bert_ner_japanese Reranker: cross_encoder_mmarco, japanese-reranker-cross-encoder, ruri-v3-reranker Sentence generation: gpt2, rinna_gpt2 Sentiment analysis: bert_sentiment_analysis, bert_tweets_sentiment Summarize: bert_sum_ext, presumm, t5_base_japanese_title_generation, t5_base_summarization Translation: fugumt-en-ja, fugumt-ja-en Zero shot classification: bert_zero_shot_classification, multilingual-minilmv2 |
| Network intrusion detection | bert-network-packet-flow-header-payload, falcon-adapter-network-packet |
 | Neural rendering | nerf, TripoSR |
| NSFW detector | clip-based-nsfw-detector |
 | Object detection | CNN: yolov1-tiny, yolov2, yolov2-tiny, maskrcnn, yolov3, yolov3-tiny, mobilenet_ssd, m2det, centernet, yolact, efficientdet, pedestrian_detection, crowd_det, yolov4, yolov4-tiny, yolov5, poly_yolo, nanodet, yolor, yolox, picodet, yolox-ti-lite, yolov7, fastest-det, yolov, yolov6, damo_yolo, yolov8, yolox_body_head_hand_face, yolov9, yolov10, yolov11, yolov12 Transformer: detr, glip, dab-detr, detic, groundingdino, rt-detr-v2 Specific target: traffic-sign-detection, sku110k-densedet, footandball, qrcode_wechatqrcode, mobile_object_localizer, layout_parsing |
 | Object detection 3d | 3d_bbox, d4lcn, egonet, mediapipe_objectron, 3d-object-detection.pytorch, did_m3d |
 | Object tracking | deepsort, person_reid_baseline_pytorch, abd_net, deepsort_vehicle, qd-3dt, centroids-reid, siam-mot, bytetrack, strong_sort, samurai |
 | Optical flow estimation | raft, cotracker3 |
 | Point segmentation | pointnet_pytorch |
 | Pose estimation | openpose, posenet, pose_resnet, lightweight-human-pose-estimation, animalpose, efficientpose, blazepose, mediapipe_holistic, movenet, ap-10k, e2pose |
 | Pose estimation 3d | pose-hg-3d, 3d-pose-baseline, lightweight-human-pose-estimation-3d, 3dmppe_posenet, gast, blazepose-fullbody, mediapipe_pose_world_landmarks |
 | Road detection | road-segmentation-adas, codes-for-lane-detection, ultra-fast-lane-detection, polylanenet, roneld, lstr, yolop, cdnet, hybridnets |
 | Rotation prediction | rotnet |
 | Style transfer | adain, pix2pixHD, beauty_gan, psgan, animeganv2, EleGANt |
 | Super resolution | srresnet, edsr, han, real-esrgan, swinir, rcan-it, Hat, SPAN |
 | Text detection | east, pixel_link, craft_pytorch |
 | Text recognition | etl, crnn.pytorch, deep-text-recognition-benchmark, easyocr, paddleocr, donut, ndlocr_text_recognition, paddleocr_v3 |
| Time-series forecasting | informer2020, timesfm, moirai, chronos2 |
 | Vehicle recognition | vehicle-attributes-recognition-barrier, vehicle-license-plate-detection-barrier |
 | Vision language model | llava, florence2, mobilevlm, llava-jp, qwen2_vl, qwen2.5_vl, qwen3_vl |
| Commercial model | acculus-pose |