8 repos
Quantized and fine-tuned variants of large language models (Llama 3.1, Phi-3, RLHFlow) optimized for inference endpoints and reinforcement learning from human feedback (RLHF) training pipelines. The cluster centers on model repositories configured for token and segment-based PPO training at 60k steps, with focus on conversational and text-generation tasks compatible with standard inference frameworks like Transformers and text-generation-inference.