| Blog | Documentation | Join Slack | Join Bi-Weekly Development Meeting | Slides |
SGLang is a fast serving framework for large language models and vision language models. It makes your interaction with models faster and more controllable by co-designing the backend runtime and frontend language. The core features include:
Install SGLang: See https://sgl-project.github.io/start/install.html
Send requests: See https://sgl-project.github.io/start/send_request.html
See https://sgl-project.github.io/backend/backend.html
See https://sgl-project.github.io/frontend/frontend.html
Learn more in our release blogs: v0.2 blog, v0.3 blog
Please cite our paper, SGLang: Efficient Execution of Structured Language Model Programs, if you find the project useful. We also learned from the design and reused code from the following projects: Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL.
14 commits
Python
98.3%
Rust
1.1%
| Blog | Documentation | Join Slack | Join Bi-Weekly Development Meeting | Slides |
SGLang is a fast serving framework for large language models and vision language models. It makes your interaction with models faster and more controllable by co-designing the backend runtime and frontend language. The core features include:
Install SGLang: See https://sgl-project.github.io/start/install.html
Send requests: See https://sgl-project.github.io/start/send_request.html
See https://sgl-project.github.io/backend/backend.html
See https://sgl-project.github.io/frontend/frontend.html
Learn more in our release blogs: v0.2 blog, v0.3 blog
Please cite our paper, SGLang: Efficient Execution of Structured Language Model Programs, if you find the project useful. We also learned from the design and reused code from the following projects: Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL.
14 commits
Python
98.3%
Rust
1.1%