langchaint is an opinionated, provider-neutral Python client for LLM applications. It provides fully typed, asynchronous APIs for generation, streaming, embeddings, tools, retries, and billing. The application owns the agent loop.
Alpha: the API may change without notice.
LLM.bind(), then call generate_one(), generate_many(), or stream_one() on the resulting BoundLLM.response_format=Answer gives generate_one() the return type Response[Answer]. Binding tools as well adds ToolCallTurn[Answer] to that return type..kind with editor autocomplete and no class imports.GenerationError values retain provider-reported usage from every recorded attempt, including billed retries.stream_one() returns an async context manager and async iterator. final() returns the typed result with its usage.langchaint requires Python 3.13 or newer.
Install the extra for each backend you use:
pip install "langchaint[openai]"
| Backend | Class | Install |
|---|---|---|
| Anthropic | Anthropic | langchaint[anthropic] |
| Anthropic on Amazon Bedrock | AnthropicBedrock | langchaint[anthropic-bedrock] |
| Cohere embeddings on Amazon Bedrock | CohereBedrock | langchaint[cohere-bedrock] |
| DeepSeek | DeepSeek | langchaint[deepseek] |
| Gemini | Gemini | langchaint[gemini] |
| OpenAI | OpenAI | langchaint[openai] |
| OpenAI embeddings | OpenAI | langchaint[openai-embedding] |
| OpenAI on Amazon Bedrock | OpenAIBedrock | langchaint[openai-bedrock] |
Install langchaint[tracing] for OpenTelemetry tracing.
import asyncio
from pydantic import BaseModel
from langchaint.openai import OpenAI
class Answer(BaseModel):
answer: str
confidence: float
async def main() -> None:
assistant = (
OpenAI()
.model("gpt-5.6-terra")
.bind(
system_prompt="Answer clearly and concisely.",
response_format=Answer,
)
)
response = await assistant.generate_one("Why is the sky blue?")
print(response.output.answer)
print(response.usage.cost_in_usd)
asyncio.run(main())
The Pydantic model validates the provider response.
generate_many() returns one result per input in input order.
A terminal failure becomes that input's GenerationError, so sibling results remain available.
Create one OpenAI for each rate-limit quota:
openai = OpenAI(
max_concurrent_requests=8,
max_request_starts_per_second=50.0,
)
fast_model = openai.model("gpt-5.6-luna")
strong_model = openai.model("gpt-5.6-sol")
A rate-limit response pauses request starts across the shared quota. After a transient failure local to one request, langchaint waits and retries that request.
text_assistant = OpenAI().model("gpt-5.6-terra").bind()
async with text_assistant.stream_one("Explain photosynthesis.") as stream:
async for item in stream:
if isinstance(item, str):
print(item, end="", flush=True)
response = await stream.final()
final() consumes the remaining stream and returns the assembled result.
The application controls turn limits, state, approvals, model changes, and persistence.
messages: list[Message] = [UserMessage(content=prompt)]
for _ in range(max_turns):
result = await bound.generate_one(messages)
match result.kind:
case "tool_call_turn":
messages.append(result.assistant_message)
outcomes = await bound.tool_manager.dispatch_many(result.tool_calls)
messages.extend(outcome.tool_message for outcome in outcomes)
case "response":
return result.output
raise RuntimeError("model did not finish within max_turns")
ToolManager.dispatch_many() runs tool calls concurrently and preserves their order.
See examples/02_tool_loop.py for a complete typed tool loop.
response.usage.cost_in_usd includes every billed retry recorded for the call.
GenerationError.usage preserves the recorded cost of failed calls.
See examples/README.md for complete examples.
langchaint uses the MIT License.
Hacker News (1)
Python
99.9%
langchaint is an opinionated, provider-neutral Python client for LLM applications. It provides fully typed, asynchronous APIs for generation, streaming, embeddings, tools, retries, and billing. The application owns the agent loop.
Alpha: the API may change without notice.
LLM.bind(), then call generate_one(), generate_many(), or stream_one() on the resulting BoundLLM.response_format=Answer gives generate_one() the return type Response[Answer]. Binding tools as well adds ToolCallTurn[Answer] to that return type..kind with editor autocomplete and no class imports.GenerationError values retain provider-reported usage from every recorded attempt, including billed retries.stream_one() returns an async context manager and async iterator. final() returns the typed result with its usage.langchaint requires Python 3.13 or newer.
Install the extra for each backend you use:
pip install "langchaint[openai]"
| Backend | Class | Install |
|---|---|---|
| Anthropic | Anthropic | langchaint[anthropic] |
| Anthropic on Amazon Bedrock | AnthropicBedrock | langchaint[anthropic-bedrock] |
| Cohere embeddings on Amazon Bedrock | CohereBedrock | langchaint[cohere-bedrock] |
| DeepSeek | DeepSeek | langchaint[deepseek] |
| Gemini | Gemini | langchaint[gemini] |
| OpenAI | OpenAI | langchaint[openai] |
| OpenAI embeddings | OpenAI | langchaint[openai-embedding] |
| OpenAI on Amazon Bedrock | OpenAIBedrock | langchaint[openai-bedrock] |
Install langchaint[tracing] for OpenTelemetry tracing.
import asyncio
from pydantic import BaseModel
from langchaint.openai import OpenAI
class Answer(BaseModel):
answer: str
confidence: float
async def main() -> None:
assistant = (
OpenAI()
.model("gpt-5.6-terra")
.bind(
system_prompt="Answer clearly and concisely.",
response_format=Answer,
)
)
response = await assistant.generate_one("Why is the sky blue?")
print(response.output.answer)
print(response.usage.cost_in_usd)
asyncio.run(main())
The Pydantic model validates the provider response.
generate_many() returns one result per input in input order.
A terminal failure becomes that input's GenerationError, so sibling results remain available.
Create one OpenAI for each rate-limit quota:
openai = OpenAI(
max_concurrent_requests=8,
max_request_starts_per_second=50.0,
)
fast_model = openai.model("gpt-5.6-luna")
strong_model = openai.model("gpt-5.6-sol")
A rate-limit response pauses request starts across the shared quota. After a transient failure local to one request, langchaint waits and retries that request.
text_assistant = OpenAI().model("gpt-5.6-terra").bind()
async with text_assistant.stream_one("Explain photosynthesis.") as stream:
async for item in stream:
if isinstance(item, str):
print(item, end="", flush=True)
response = await stream.final()
final() consumes the remaining stream and returns the assembled result.
The application controls turn limits, state, approvals, model changes, and persistence.
messages: list[Message] = [UserMessage(content=prompt)]
for _ in range(max_turns):
result = await bound.generate_one(messages)
match result.kind:
case "tool_call_turn":
messages.append(result.assistant_message)
outcomes = await bound.tool_manager.dispatch_many(result.tool_calls)
messages.extend(outcome.tool_message for outcome in outcomes)
case "response":
return result.output
raise RuntimeError("model did not finish within max_turns")
ToolManager.dispatch_many() runs tool calls concurrently and preserves their order.
See examples/02_tool_loop.py for a complete typed tool loop.
response.usage.cost_in_usd includes every billed retry recorded for the call.
GenerationError.usage preserves the recorded cost of failed calls.
See examples/README.md for complete examples.
langchaint uses the MIT License.
Hacker News (1)
Python
99.9%