LLM Inference & RLHF Training

8 repos

Quantized and fine-tuned variants of large language models (Llama 3.1, Phi-3, RLHFlow) optimized for inference endpoints and reinforcement learning from human feedback (RLHF) training pipelines. The cluster centers on model repositories configured for token and segment-based PPO training at 60k steps, with focus on conversational and text-generation tasks compatible with standard inference frameworks like Transformers and text-generation-inference.

Python · 1
conversational ·1
text-generation-inference ·1
transformers ·1
endpoints_compatible ·1
llama ·1
safetensors ·1
text-generation ·1
custom_code ·0
phi3 ·0