abubakarsiddik31/golem

Go-first framework for dependable AI agents: typed dependencies and outputs, explicit tools, composable models, observable runs

Go

3

169 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a zero-dependency Go framework for running agent loops with local models (Ollama, LM Studio, vLLM) (r/LocalLLM)

Most AI agent libraries (LangChain, PydanticAI, CrewAI) live in Python. That works well enough when you are hacking on a prototype, but if you want to deploy a small daemon or backend service that talks to a local model via Ollama or LM Studio, dragging along a Python runtime with hundreds of…

2

Oct 6, 2026

Show HN: Golem – Zero-dependency, type-safe AI agent framework in pure Go

2

Oct 6, 2026

README

Golem: a type-safe AI agent framework for Go

A Go-first framework for building dependable, production-grade AI agents.
Compile-time type safety, zero external dependencies, explicit execution control, and deep MCP integration.

Go Reference CI Status Documentation Zero Dependencies License: MIT Release v0.8.5

Read the Documentation »

Why Golem? • Quick Start • Supported Providers • Architecture • Common Tools & MCP • Testing • Guides Index • Examples • Contributing


Why Golem?

Python frameworks like Pydantic AI made agent prototyping accessible with typed schemas and clean ergonomics. However, deploying AI agents to mission-critical production systems demands the strengths of Go: high-throughput concurrency, predictable memory usage, fast execution, small deployment binaries, and compile-time correctness.

Golem bridges that gap by offering idiomatic, enterprise-ready agent abstractions designed specifically for Go engineers:

  • 🛡️ Compile-Time Type Safety & Generics: Agents are declared as Agent[Deps, Output]. Tools access strongly-typed dependencies (databases, auth contexts, HTTP clients) via RunContext[Deps]. No map[string]any spaghetti or unexpected runtime reflection errors.
  • 📦 Zero External Dependencies: Built exclusively on Go's standard library. Instant compilation, tiny container images, and a clean security profile free of supply-chain vulnerabilities.
  • 🌐 Provider Agnostic: Native adapters for OpenAI, Anthropic (with prompt caching and thinking signatures), Google Gemini, AWS Bedrock (Converse API with SigV4), Azure OpenAI, and local offline models with Ollama and LM Studio.
  • 🔌 Native Model Context Protocol (MCP): Full-featured MCP client supporting both stdio and streaming HTTP transports, effortlessly turning external MCP servers into typed agent tools.
  • 🔄 Production Resilience & Self-Correction: In-loop model self-correction (ModelRetry), automated multi-model fallbacks, exponential backoff, per-tool deadlines, and clean in-tool cancellation (tool.Canceled).
  • 📊 Auditable & Observable Evidence: Every run produces normalized messages, durable additive JSON, live streamable run events, reasoning/thinking token capture, and RunError.Partial—preserving all intermediate tool results even when a run fails or gets cancelled.
  • ⏸️ Human-in-the-Loop & Deferred Execution: Pause agent runs cleanly when tools require human sign-off or external async triggers, and resume deterministically with full preserved state.
  • 🛠️ Batteries-Included Tooling: Layout-aware PDF extraction (pdfextract), multi-format document extraction (docextract: Word, Excel, PowerPoint, Markdown, CSV), SSR web reader (webfetch), workspace file accessor (fileread), shell command execution (shell), and Agent Skills progressive loading (skills).
  • 💰 Budget & Cost Guards: Pre-send token estimation, token-budgeted history truncation, client-side untrusted history sanitization (SanitizeHistory), and user-defined price tables to calculate and cap dollar costs per run.

Architecture

Golem orchestrates agents through a transparent, observable execution loop:

                         ┌────────────────────────────────────────┐
                         │         golem.RunContext[Deps]         │
                         │        (Typed Run Dependencies)        │
                         └───────────────────┬────────────────────┘
                                             ▼
┌─────────────────┐          ┌───────────────────────────────┐          ┌─────────────────┐
│     Prompt      │ ───────► │     golem.Agent[Deps, Out]    │ ───────► │  Typed Output   │
│ (Text, Images,  │          │                               │          │  Result[Output] │
│ Docs, Audio)    │          │  ┌─────────────────────────┐  │          └─────────────────┘
└─────────────────┘          │  │ Observable Loop         │  │                   │
                             │  │ - Self-Correction       │  │                   ▼
                             │  │ - Retries & Fallbacks   │  │          ┌─────────────────┐
                             │  │ - Token & Cost Bounds   │  │          │ Durable Evidence│
                             │  │ - Run Events Stream     │  │          │ (Messages, Cost,│
                             │  └────────────┬────────────┘  │          │  Token Usage)   │
                             └───────────────┼───────────────┘          └─────────────────┘
                                             │
                      ┌──────────────────────┴──────────────────────┐
                      ▼                                             ▼
        ┌───────────────────────────┐                 ┌───────────────────────────┐
        │      Model Adapters       │                 │      Tools & Protocols    │
        │  • OpenAI & Azure OpenAI  │                 │  • Strongly Typed Tools   │
        │  • Anthropic Claude       │                 │  • MCP Client (stdio/HTTP)│
        │  • Google Gemini          │                 │  • PDF & Doc Extractors   │
        │  • AWS Bedrock (SigV4)    │                 │  • Web Fetch & Shell      │
        │  • Local (Ollama/LMStudio)│                 │  • Agent Skills (SKILL.md)│
        └───────────────────────────┘                 └───────────────────────────┘

Installation

go get github.com/abubakarsiddik31/golem

Requires Go 1.26.5 or newer. Zero external dependencies.


Quick Start

1. Minimal Agent

Initialize a model adapter, create a typed agent, and execute a prompt:

package main

import (
	"context"
	"fmt"
	"os"

	"github.com/abubakarsiddik31/golem"
	"github.com/abubakarsiddik31/golem/model"
	"github.com/abubakarsiddik31/golem/providers/openai"
)

func main() {
	client, err := openai.New(openai.Config{
		APIKey: os.Getenv("OPENAI_API_KEY"),
		Model:  "gpt-4o-mini",
	})
	if err != nil {
		panic(err)
	}

	agent, err := golem.New[struct{}, string](client,
		golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
			return r.Message.Content, nil
		}),
	)
	if err != nil {
		panic(err)
	}

	result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "Why is Go ideal for AI agents?")
	if err != nil {
		panic(err)
	}

	fmt.Println(result.Output)
	fmt.Printf("Tokens: %d input, %d output\n", result.Usage.InputTokens, result.Usage.OutputTokens)
}

2. Typed Tools with Dependency Injection

Golem tools are strongly typed and receive dependencies via RunContext[Deps]. No global state, no untyped maps:

type Database struct {
	Users map[int]string
}

// Declare a tool typed to Database dependencies
getUser := tool.MustNew(tool.Tool[Database]{
	Name:        "get_user",
	Description: "Look up a user name by their ID.",
	Schema: json.RawMessage(`{
		"type": "object",
		"properties": {"id": {"type": "integer"}},
		"required": ["id"]
	}`),
	Exec: func(ctx context.Context, db Database, args json.RawMessage) (tool.Result, error) {
		var input struct {
			ID int `json:"id"`
		}
		if err := json.Unmarshal(args, &input); err != nil {
			return tool.Result{}, err
		}
		name, ok := db.Users[input.ID]
		if !ok {
			return tool.Text("User not found"), nil
		}
		return tool.Text(name), nil
	},
})

// Create an agent parameterized with Database dependencies
agent, err := golem.New[Database, string](client,
	golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
		return r.Message.Content, nil
	}),
	golem.WithTools[Database, string](getUser),
)

// Run passing the typed dependency instance
db := Database{Users: map[int]string{42: "Alice"}}
result, err := agent.Run(ctx, golem.RunContext[Database]{Deps: db}, "Who is user 42?")

3. Structured Output & Schema Validation

Guarantee your agent returns strongly-typed Go structs with golem.DecodeJSON[T]() and strict schema enforcement:

type WeatherReport struct {
	City        string  `json:"city"`
	Temperature float64 `json:"temperature_celsius"`
	Condition   string  `json:"condition"`
}

agent, err := golem.New[struct{}, WeatherReport](client,
	golem.DecodeJSON[WeatherReport](),
	golem.WithOutputSchema[struct{}, WeatherReport](json.RawMessage(`{
		"type": "object",
		"properties": {
			"city": {"type": "string"},
			"temperature_celsius": {"type": "number"},
			"condition": {"type": "string"}
		},
		"required": ["city", "temperature_celsius", "condition"],
		"additionalProperties": false
	}`)),
)

result, err := agent.Run(ctx, golem.RunContext[struct{}]{}, "Forecast for Lagos, Nigeria.")
fmt.Printf("%s: %.1f°C (%s)\n", result.Output.City, result.Output.Temperature, result.Output.Condition)

Supported Providers

Golem ships with standard-library-only adapters for all major frontier and open-weight models:

ProviderAdapter PackageStreamingThinking / ReasoningMultimodalEmbeddingsToken Counting
OpenAIproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
Anthropicproviders/anthropic:white_check_mark::white_check_mark::white_check_mark:—:white_check_mark:
Google Geminiproviders/gemini:white_check_mark::white_check_mark::white_check_mark::white_check_mark::white_check_mark:
AWS Bedrockproviders/bedrock:white_check_mark::white_check_mark::white_check_mark:—:white_check_mark:
Azure OpenAIproviders/azure:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
Ollama / Localproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
OpenAI-Compatibleproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—

Batteries-Included Tools

Golem includes pre-built common tools written entirely in pure Go:

PackageCapabilityFeatures
mcpModel Context Protocol ClientConnects to any MCP tool server over stdio or streaming HTTP (SSE).
pdfextractLayout-Aware PDF ExtractionHigh-performance PDF parser preserving reading order, layout, tables, and embedded images.
docextractMulti-Format Document ExtractionExtracts text and structure from Word (.docx), Excel (.xlsx), PowerPoint (.pptx), CSV, and Markdown.
webfetchWeb ExtractionFetches URLs and returns clean, agent-readable text without browser overhead.
filereadFile ReaderSafe workspace file reading with path boundaries and clean formatting.
shellCommand ExecutionIsolated command execution with timeout handling and combined stdout/stderr output.
skillsAgent Skills LoaderDiscovers and loads standard SKILL.md skill folders on demand for progressive prompt enrichment.

Testing Without a Provider

Never mock HTTP endpoints or pay for tokens in unit tests. Golem includes testmodel, a fully deterministic, offline model implementation:

package main

import (
	"context"
	"testing"

	"github.com/abubakarsiddik31/golem"
	"github.com/abubakarsiddik31/golem/model"
	"github.com/abubakarsiddik31/golem/testmodel"
)

func TestAgent(t *testing.T) {
	client := testmodel.New().Respond(
		model.Response{Message: model.Message{Role: model.RoleAssistant, Content: "pong"}},
	)

	agent, _ := golem.New[struct{}, string](client,
		golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
			return r.Message.Content, nil
		}),
	)

	result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "ping")
	if err != nil || result.Output != "pong" {
		t.Fatalf("unexpected result: %v, output: %s", err, result.Output)
	}
}

Documentation

The feature guides provide complete references for each capability and mirror the published documentation site at abubakarsiddik31.github.io/golem.

GuideCovers
Getting startedThe smallest agent, result shape, error stages
ProvidersOpenAI-compatible and Anthropic adapters, error classification
EmbeddingsThe embedding.Embedder port: queries, documents, usage
Token countingThe tokens.Counter port: budgets, pre-send limits
CostUser-supplied pricing: Result.Cost and cost bounds
Tools and dependenciesTyped tools, dependencies, and controlled parallel execution
Web fetchThe webfetch common tool: URLs as agent-readable text
File readThe fileread common tool: workspace files as agent-readable text
Command executionThe shell common tool: one command, combined output
PDF extractThe pdfextract common tool: PDF documents as structured Markdown with tables and images
Document extractThe docextract common tool: Word, Excel, PowerPoint, Markdown, CSV, and multi-format documents
Agent skillsThe skills common tool: standard SKILL.md folders loaded on demand
MCP clientBridging Model Context Protocol servers into agent tools
Agent delegationOne agent exposed as another agent's tool
Tool timeoutsContext-aware deadlines for individual tool calls
Conversations and historyMulti-turn runs, durable message JSON, history trimming
Multimodal inputImages, documents, audio, and video in prompts, per-provider mapping
Structured outputOutput schemas, tool-mode output, DecodeJSON
Self-correctionOutput and tool rejection budgets (ModelRetry)
RetriesTransient model failures, backoff, fallback models
StreamingRunStream, the streaming capability port, SSE adapters
Run eventsObserving attempts, tool calls, and corrections as they happen
ThinkingReasoning models: requesting thinking, keeping signatures, replay
Usage limitsBounding tokens, requests, and tool calls
Testing without a providerDeterministic fakes, contract assertions
Deferred toolsApprovals and external results: pausing a run and resuming it

Design records live in docs/adr/; each guide links the ADR that decided its behavior.


Examples

Runnable programs live in examples/; provider-backed examples read their API key from the environment and exit cleanly with instructions when unset.

ExampleShows
minimalSmallest agent against an OpenAI-compatible API
toolsTyped tool with a run dependency
structured-outputOutput schema + JSON decoding
structured-output-toolTool-mode structured output
conversationInteractive multi-turn chat with history
streamingRunStream printing fragments as they arrive
run-eventsWithRunEvents printing the event sequence of a run
mcp-clientMCP server bridged into agent tools over stdio
mcp-httpMCP server bridged over streamable HTTP
skillsThe skills common tool loading a standard SKILL.md folder
web-fetchThe webfetch common tool fetching a local test page
file-readThe fileread common tool reading a workspace file
command-executionThe shell common tool running one local command
pdf-extractThe pdfextract common tool extracting tables and reading order, offline
doc-extractThe docextract common tool extracting Word, Excel, and Markdown, offline
delegationA specialist agent delegated to as a tool
tool-resultsTools returning parts and definitive failures, offline
self-correctionTool rejecting correctable arguments
fallbackPrimary model with a fallback and a request bound
token-countingPre-send limits and budget-bounded history over the tokens.Counter port
costUser-supplied pricing: Result.Cost and cost-bounded runs, offline
embeddingsSemantic search over the embedding.Embedder port
multimodal-inputPrompts with images and document attachments
thinkingAdaptive thinking with reasoning blocks and signatures
run-cancellationA tool ending the run deliberately with tool.Canceled, resuming evidence, offline
run-idsRun and conversation identity across chained and forked runs, offline
deferred-toolsPausing runs for approvals or external results, and resuming, offline
partial-evidenceA failed run's RunError.Partial evidence resumed with history, offline
history-repairNormalizing a damaged conversation with a report, offline
history-sanitizationSanitizing a client-submitted history at the trust boundary, offline
anthropicAnthropic Messages API adapter
geminiGoogle Gemini GenerateContent adapter
azureAzure OpenAI deployment adapter
bedrockAWS Bedrock Converse adapter with SigV4
local-modelsOllama or LM Studio through the OpenAI-compatible adapter
testing-without-a-providerScripted fake model, offline and deterministic

To run any example:

# Run with OpenAI
OPENAI_API_KEY=sk-... go run ./examples/minimal

# Run locally with Ollama or LM Studio
GOLEM_LOCAL_BASE_URL=http://localhost:11434/v1 go run ./examples/local-models

# Run offline examples (no credentials needed)
go run ./examples/testing-without-a-provider
go run ./examples/pdf-extract
go run ./examples/deferred-tools
go run ./examples/partial-evidence

Package Structure

golem/        Agent configuration and typed run API
model/        Provider-neutral model request/response contract
tool/         Tool declarations and execution contracts
mcp/          Model Context Protocol client (stdio & streaming HTTP)
pdfextract/   Common tool: layout-aware PDF parser (tables, images, text)
docextract/   Common tool: Word, Excel, PowerPoint, CSV, and Markdown extractor
webfetch/     Common tool: fetch a URL as agent-readable text
fileread/     Common tool: read a file as agent-readable text
shell/        Common tool: run one command, return combined output
skills/       Common tool: standard SKILL.md folder progressive loader
providers/    Stdlib-only adapters implementing model.Model
testmodel/    Deterministic in-memory model doubles for unit testing
internal/     Execution runner loop and private mechanics
examples/     Runnable programs per capability
docs/guides/  Feature guides (source of truth for behavior)
docs/adr/     Decisions that shape public contracts

Status & Roadmap

Golem is currently at v0.8.5.

The core execution contract is frozen and verified with continuous race-detector CI, memory fuzzing, and deterministic offline tests. The public API adheres strictly to additive-only changes on the road to v1.0.0.

  • Resilient Execution: Self-correction loops (ModelRetry), fallback models, exponential backoff, per-tool timeouts, in-tool cancellation (tool.Canceled), and partial evidence preservation (RunError.Partial).
  • Comprehensive Multimodal: Text, images, PDF documents, audio, and video inputs mapped natively across all model adapters.
  • Observability & Durability: Normalized conversation messages, durable additive JSON serialization, reasoning/thinking token capture with provider signatures, and streamable run events.
  • Production Guardrails: Pre-send token estimation, token-budgeted history truncation, client-side untrusted history sanitization (SanitizeHistory), history repair (NormalizeHistory), and user-configurable cost/token bounding.
  • Rich Tool Ecosystem: First-class Model Context Protocol (MCP) client over stdio and HTTP, layout-aware PDF extraction, multi-format doc parsing, sandboxed shell execution, SSR web fetch, and progressive Agent Skills loading.
  • Zero Dependencies: Built exclusively on the Go standard library.

For architecture rationale, read the foundation brief. For upcoming milestones, read the development roadmap.


Development

Run the test suite and verification checks:

go test ./...
go test -race ./...
go vet ./...

The feature guides publish as the official documentation site; preview it locally with mkdocs serve (see docs/website.md). Brand assets and guidelines live in assets/brand/.


Community & Contributing

Golem is an open-source project and actively welcomes community contributions!


License

Released under the MIT License.

agent-framework
ai-agents
anthropic
gemini
go
llm
openai

abubakarsiddik31/golem

Go-first framework for dependable AI agents: typed dependencies and outputs, explicit tools, composable models, observable runs

Go

3

169 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a zero-dependency Go framework for running agent loops with local models (Ollama, LM Studio, vLLM) (r/LocalLLM)

Most AI agent libraries (LangChain, PydanticAI, CrewAI) live in Python. That works well enough when you are hacking on a prototype, but if you want to deploy a small daemon or backend service that talks to a local model via Ollama or LM Studio, dragging along a Python runtime with hundreds of…

2

Oct 6, 2026

Show HN: Golem – Zero-dependency, type-safe AI agent framework in pure Go

2

Oct 6, 2026

README

Golem: a type-safe AI agent framework for Go

A Go-first framework for building dependable, production-grade AI agents.
Compile-time type safety, zero external dependencies, explicit execution control, and deep MCP integration.

Go Reference CI Status Documentation Zero Dependencies License: MIT Release v0.8.5

Read the Documentation »

Why Golem? • Quick Start • Supported Providers • Architecture • Common Tools & MCP • Testing • Guides Index • Examples • Contributing


Why Golem?

Python frameworks like Pydantic AI made agent prototyping accessible with typed schemas and clean ergonomics. However, deploying AI agents to mission-critical production systems demands the strengths of Go: high-throughput concurrency, predictable memory usage, fast execution, small deployment binaries, and compile-time correctness.

Golem bridges that gap by offering idiomatic, enterprise-ready agent abstractions designed specifically for Go engineers:

  • 🛡️ Compile-Time Type Safety & Generics: Agents are declared as Agent[Deps, Output]. Tools access strongly-typed dependencies (databases, auth contexts, HTTP clients) via RunContext[Deps]. No map[string]any spaghetti or unexpected runtime reflection errors.
  • 📦 Zero External Dependencies: Built exclusively on Go's standard library. Instant compilation, tiny container images, and a clean security profile free of supply-chain vulnerabilities.
  • 🌐 Provider Agnostic: Native adapters for OpenAI, Anthropic (with prompt caching and thinking signatures), Google Gemini, AWS Bedrock (Converse API with SigV4), Azure OpenAI, and local offline models with Ollama and LM Studio.
  • 🔌 Native Model Context Protocol (MCP): Full-featured MCP client supporting both stdio and streaming HTTP transports, effortlessly turning external MCP servers into typed agent tools.
  • 🔄 Production Resilience & Self-Correction: In-loop model self-correction (ModelRetry), automated multi-model fallbacks, exponential backoff, per-tool deadlines, and clean in-tool cancellation (tool.Canceled).
  • 📊 Auditable & Observable Evidence: Every run produces normalized messages, durable additive JSON, live streamable run events, reasoning/thinking token capture, and RunError.Partial—preserving all intermediate tool results even when a run fails or gets cancelled.
  • ⏸️ Human-in-the-Loop & Deferred Execution: Pause agent runs cleanly when tools require human sign-off or external async triggers, and resume deterministically with full preserved state.
  • 🛠️ Batteries-Included Tooling: Layout-aware PDF extraction (pdfextract), multi-format document extraction (docextract: Word, Excel, PowerPoint, Markdown, CSV), SSR web reader (webfetch), workspace file accessor (fileread), shell command execution (shell), and Agent Skills progressive loading (skills).
  • 💰 Budget & Cost Guards: Pre-send token estimation, token-budgeted history truncation, client-side untrusted history sanitization (SanitizeHistory), and user-defined price tables to calculate and cap dollar costs per run.

Architecture

Golem orchestrates agents through a transparent, observable execution loop:

                         ┌────────────────────────────────────────┐
                         │         golem.RunContext[Deps]         │
                         │        (Typed Run Dependencies)        │
                         └───────────────────┬────────────────────┘
                                             ▼
┌─────────────────┐          ┌───────────────────────────────┐          ┌─────────────────┐
│     Prompt      │ ───────► │     golem.Agent[Deps, Out]    │ ───────► │  Typed Output   │
│ (Text, Images,  │          │                               │          │  Result[Output] │
│ Docs, Audio)    │          │  ┌─────────────────────────┐  │          └─────────────────┘
└─────────────────┘          │  │ Observable Loop         │  │                   │
                             │  │ - Self-Correction       │  │                   ▼
                             │  │ - Retries & Fallbacks   │  │          ┌─────────────────┐
                             │  │ - Token & Cost Bounds   │  │          │ Durable Evidence│
                             │  │ - Run Events Stream     │  │          │ (Messages, Cost,│
                             │  └────────────┬────────────┘  │          │  Token Usage)   │
                             └───────────────┼───────────────┘          └─────────────────┘
                                             │
                      ┌──────────────────────┴──────────────────────┐
                      ▼                                             ▼
        ┌───────────────────────────┐                 ┌───────────────────────────┐
        │      Model Adapters       │                 │      Tools & Protocols    │
        │  • OpenAI & Azure OpenAI  │                 │  • Strongly Typed Tools   │
        │  • Anthropic Claude       │                 │  • MCP Client (stdio/HTTP)│
        │  • Google Gemini          │                 │  • PDF & Doc Extractors   │
        │  • AWS Bedrock (SigV4)    │                 │  • Web Fetch & Shell      │
        │  • Local (Ollama/LMStudio)│                 │  • Agent Skills (SKILL.md)│
        └───────────────────────────┘                 └───────────────────────────┘

Installation

go get github.com/abubakarsiddik31/golem

Requires Go 1.26.5 or newer. Zero external dependencies.


Quick Start

1. Minimal Agent

Initialize a model adapter, create a typed agent, and execute a prompt:

package main

import (
	"context"
	"fmt"
	"os"

	"github.com/abubakarsiddik31/golem"
	"github.com/abubakarsiddik31/golem/model"
	"github.com/abubakarsiddik31/golem/providers/openai"
)

func main() {
	client, err := openai.New(openai.Config{
		APIKey: os.Getenv("OPENAI_API_KEY"),
		Model:  "gpt-4o-mini",
	})
	if err != nil {
		panic(err)
	}

	agent, err := golem.New[struct{}, string](client,
		golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
			return r.Message.Content, nil
		}),
	)
	if err != nil {
		panic(err)
	}

	result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "Why is Go ideal for AI agents?")
	if err != nil {
		panic(err)
	}

	fmt.Println(result.Output)
	fmt.Printf("Tokens: %d input, %d output\n", result.Usage.InputTokens, result.Usage.OutputTokens)
}

2. Typed Tools with Dependency Injection

Golem tools are strongly typed and receive dependencies via RunContext[Deps]. No global state, no untyped maps:

type Database struct {
	Users map[int]string
}

// Declare a tool typed to Database dependencies
getUser := tool.MustNew(tool.Tool[Database]{
	Name:        "get_user",
	Description: "Look up a user name by their ID.",
	Schema: json.RawMessage(`{
		"type": "object",
		"properties": {"id": {"type": "integer"}},
		"required": ["id"]
	}`),
	Exec: func(ctx context.Context, db Database, args json.RawMessage) (tool.Result, error) {
		var input struct {
			ID int `json:"id"`
		}
		if err := json.Unmarshal(args, &input); err != nil {
			return tool.Result{}, err
		}
		name, ok := db.Users[input.ID]
		if !ok {
			return tool.Text("User not found"), nil
		}
		return tool.Text(name), nil
	},
})

// Create an agent parameterized with Database dependencies
agent, err := golem.New[Database, string](client,
	golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
		return r.Message.Content, nil
	}),
	golem.WithTools[Database, string](getUser),
)

// Run passing the typed dependency instance
db := Database{Users: map[int]string{42: "Alice"}}
result, err := agent.Run(ctx, golem.RunContext[Database]{Deps: db}, "Who is user 42?")

3. Structured Output & Schema Validation

Guarantee your agent returns strongly-typed Go structs with golem.DecodeJSON[T]() and strict schema enforcement:

type WeatherReport struct {
	City        string  `json:"city"`
	Temperature float64 `json:"temperature_celsius"`
	Condition   string  `json:"condition"`
}

agent, err := golem.New[struct{}, WeatherReport](client,
	golem.DecodeJSON[WeatherReport](),
	golem.WithOutputSchema[struct{}, WeatherReport](json.RawMessage(`{
		"type": "object",
		"properties": {
			"city": {"type": "string"},
			"temperature_celsius": {"type": "number"},
			"condition": {"type": "string"}
		},
		"required": ["city", "temperature_celsius", "condition"],
		"additionalProperties": false
	}`)),
)

result, err := agent.Run(ctx, golem.RunContext[struct{}]{}, "Forecast for Lagos, Nigeria.")
fmt.Printf("%s: %.1f°C (%s)\n", result.Output.City, result.Output.Temperature, result.Output.Condition)

Supported Providers

Golem ships with standard-library-only adapters for all major frontier and open-weight models:

ProviderAdapter PackageStreamingThinking / ReasoningMultimodalEmbeddingsToken Counting
OpenAIproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
Anthropicproviders/anthropic:white_check_mark::white_check_mark::white_check_mark:—:white_check_mark:
Google Geminiproviders/gemini:white_check_mark::white_check_mark::white_check_mark::white_check_mark::white_check_mark:
AWS Bedrockproviders/bedrock:white_check_mark::white_check_mark::white_check_mark:—:white_check_mark:
Azure OpenAIproviders/azure:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
Ollama / Localproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—
OpenAI-Compatibleproviders/openai:white_check_mark::white_check_mark::white_check_mark::white_check_mark:—

Batteries-Included Tools

Golem includes pre-built common tools written entirely in pure Go:

PackageCapabilityFeatures
mcpModel Context Protocol ClientConnects to any MCP tool server over stdio or streaming HTTP (SSE).
pdfextractLayout-Aware PDF ExtractionHigh-performance PDF parser preserving reading order, layout, tables, and embedded images.
docextractMulti-Format Document ExtractionExtracts text and structure from Word (.docx), Excel (.xlsx), PowerPoint (.pptx), CSV, and Markdown.
webfetchWeb ExtractionFetches URLs and returns clean, agent-readable text without browser overhead.
filereadFile ReaderSafe workspace file reading with path boundaries and clean formatting.
shellCommand ExecutionIsolated command execution with timeout handling and combined stdout/stderr output.
skillsAgent Skills LoaderDiscovers and loads standard SKILL.md skill folders on demand for progressive prompt enrichment.

Testing Without a Provider

Never mock HTTP endpoints or pay for tokens in unit tests. Golem includes testmodel, a fully deterministic, offline model implementation:

package main

import (
	"context"
	"testing"

	"github.com/abubakarsiddik31/golem"
	"github.com/abubakarsiddik31/golem/model"
	"github.com/abubakarsiddik31/golem/testmodel"
)

func TestAgent(t *testing.T) {
	client := testmodel.New().Respond(
		model.Response{Message: model.Message{Role: model.RoleAssistant, Content: "pong"}},
	)

	agent, _ := golem.New[struct{}, string](client,
		golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
			return r.Message.Content, nil
		}),
	)

	result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "ping")
	if err != nil || result.Output != "pong" {
		t.Fatalf("unexpected result: %v, output: %s", err, result.Output)
	}
}

Documentation

The feature guides provide complete references for each capability and mirror the published documentation site at abubakarsiddik31.github.io/golem.

GuideCovers
Getting startedThe smallest agent, result shape, error stages
ProvidersOpenAI-compatible and Anthropic adapters, error classification
EmbeddingsThe embedding.Embedder port: queries, documents, usage
Token countingThe tokens.Counter port: budgets, pre-send limits
CostUser-supplied pricing: Result.Cost and cost bounds
Tools and dependenciesTyped tools, dependencies, and controlled parallel execution
Web fetchThe webfetch common tool: URLs as agent-readable text
File readThe fileread common tool: workspace files as agent-readable text
Command executionThe shell common tool: one command, combined output
PDF extractThe pdfextract common tool: PDF documents as structured Markdown with tables and images
Document extractThe docextract common tool: Word, Excel, PowerPoint, Markdown, CSV, and multi-format documents
Agent skillsThe skills common tool: standard SKILL.md folders loaded on demand
MCP clientBridging Model Context Protocol servers into agent tools
Agent delegationOne agent exposed as another agent's tool
Tool timeoutsContext-aware deadlines for individual tool calls
Conversations and historyMulti-turn runs, durable message JSON, history trimming
Multimodal inputImages, documents, audio, and video in prompts, per-provider mapping
Structured outputOutput schemas, tool-mode output, DecodeJSON
Self-correctionOutput and tool rejection budgets (ModelRetry)
RetriesTransient model failures, backoff, fallback models
StreamingRunStream, the streaming capability port, SSE adapters
Run eventsObserving attempts, tool calls, and corrections as they happen
ThinkingReasoning models: requesting thinking, keeping signatures, replay
Usage limitsBounding tokens, requests, and tool calls
Testing without a providerDeterministic fakes, contract assertions
Deferred toolsApprovals and external results: pausing a run and resuming it

Design records live in docs/adr/; each guide links the ADR that decided its behavior.


Examples

Runnable programs live in examples/; provider-backed examples read their API key from the environment and exit cleanly with instructions when unset.

ExampleShows
minimalSmallest agent against an OpenAI-compatible API
toolsTyped tool with a run dependency
structured-outputOutput schema + JSON decoding
structured-output-toolTool-mode structured output
conversationInteractive multi-turn chat with history
streamingRunStream printing fragments as they arrive
run-eventsWithRunEvents printing the event sequence of a run
mcp-clientMCP server bridged into agent tools over stdio
mcp-httpMCP server bridged over streamable HTTP
skillsThe skills common tool loading a standard SKILL.md folder
web-fetchThe webfetch common tool fetching a local test page
file-readThe fileread common tool reading a workspace file
command-executionThe shell common tool running one local command
pdf-extractThe pdfextract common tool extracting tables and reading order, offline
doc-extractThe docextract common tool extracting Word, Excel, and Markdown, offline
delegationA specialist agent delegated to as a tool
tool-resultsTools returning parts and definitive failures, offline
self-correctionTool rejecting correctable arguments
fallbackPrimary model with a fallback and a request bound
token-countingPre-send limits and budget-bounded history over the tokens.Counter port
costUser-supplied pricing: Result.Cost and cost-bounded runs, offline
embeddingsSemantic search over the embedding.Embedder port
multimodal-inputPrompts with images and document attachments
thinkingAdaptive thinking with reasoning blocks and signatures
run-cancellationA tool ending the run deliberately with tool.Canceled, resuming evidence, offline
run-idsRun and conversation identity across chained and forked runs, offline
deferred-toolsPausing runs for approvals or external results, and resuming, offline
partial-evidenceA failed run's RunError.Partial evidence resumed with history, offline
history-repairNormalizing a damaged conversation with a report, offline
history-sanitizationSanitizing a client-submitted history at the trust boundary, offline
anthropicAnthropic Messages API adapter
geminiGoogle Gemini GenerateContent adapter
azureAzure OpenAI deployment adapter
bedrockAWS Bedrock Converse adapter with SigV4
local-modelsOllama or LM Studio through the OpenAI-compatible adapter
testing-without-a-providerScripted fake model, offline and deterministic

To run any example:

# Run with OpenAI
OPENAI_API_KEY=sk-... go run ./examples/minimal

# Run locally with Ollama or LM Studio
GOLEM_LOCAL_BASE_URL=http://localhost:11434/v1 go run ./examples/local-models

# Run offline examples (no credentials needed)
go run ./examples/testing-without-a-provider
go run ./examples/pdf-extract
go run ./examples/deferred-tools
go run ./examples/partial-evidence

Package Structure

golem/        Agent configuration and typed run API
model/        Provider-neutral model request/response contract
tool/         Tool declarations and execution contracts
mcp/          Model Context Protocol client (stdio & streaming HTTP)
pdfextract/   Common tool: layout-aware PDF parser (tables, images, text)
docextract/   Common tool: Word, Excel, PowerPoint, CSV, and Markdown extractor
webfetch/     Common tool: fetch a URL as agent-readable text
fileread/     Common tool: read a file as agent-readable text
shell/        Common tool: run one command, return combined output
skills/       Common tool: standard SKILL.md folder progressive loader
providers/    Stdlib-only adapters implementing model.Model
testmodel/    Deterministic in-memory model doubles for unit testing
internal/     Execution runner loop and private mechanics
examples/     Runnable programs per capability
docs/guides/  Feature guides (source of truth for behavior)
docs/adr/     Decisions that shape public contracts

Status & Roadmap

Golem is currently at v0.8.5.

The core execution contract is frozen and verified with continuous race-detector CI, memory fuzzing, and deterministic offline tests. The public API adheres strictly to additive-only changes on the road to v1.0.0.

  • Resilient Execution: Self-correction loops (ModelRetry), fallback models, exponential backoff, per-tool timeouts, in-tool cancellation (tool.Canceled), and partial evidence preservation (RunError.Partial).
  • Comprehensive Multimodal: Text, images, PDF documents, audio, and video inputs mapped natively across all model adapters.
  • Observability & Durability: Normalized conversation messages, durable additive JSON serialization, reasoning/thinking token capture with provider signatures, and streamable run events.
  • Production Guardrails: Pre-send token estimation, token-budgeted history truncation, client-side untrusted history sanitization (SanitizeHistory), history repair (NormalizeHistory), and user-configurable cost/token bounding.
  • Rich Tool Ecosystem: First-class Model Context Protocol (MCP) client over stdio and HTTP, layout-aware PDF extraction, multi-format doc parsing, sandboxed shell execution, SSR web fetch, and progressive Agent Skills loading.
  • Zero Dependencies: Built exclusively on the Go standard library.

For architecture rationale, read the foundation brief. For upcoming milestones, read the development roadmap.


Development

Run the test suite and verification checks:

go test ./...
go test -race ./...
go vet ./...

The feature guides publish as the official documentation site; preview it locally with mkdocs serve (see docs/website.md). Brand assets and guidelines live in assets/brand/.


Community & Contributing

Golem is an open-source project and actively welcomes community contributions!


License

Released under the MIT License.

agent-framework
ai-agents
anthropic
gemini
go
llm
openai