Text Embeddings and Information Retrieval

19 repos

Tools and models for converting text into dense vector representations and retrieving relevant documents or passages. The cluster centers on LLM-based embedding models—particularly fine-tuned variants of Mistral, Llama, and other large language models—trained with supervised and unsupervised approaches for semantic similarity tasks. Repositories here focus on text classification, embeddings generation, and retrieval augmentation, serving as building blocks for search, recommendation, and semantic matching systems.

Python · 6
Jupyter Notebook · 1
language-model ·6,314
information-retrieval ·5,902
text-embedding ·5,008
embeddings ·4,371
text-classification ·4,319
text-clustering ·4,142
text-reranking ·4,136
text-semantic-similarity ·4,136
text-evaluation ·4,136
prompt-retrieval ·4,048