FastAPI → GraphQL → MCP: your whole API as ONE constant tool set composed queries, field selection, per-caller auth. Zero Decorators
Python
2
65 commits
updated Oct 5, 2026
Turn any FastAPI router into a GraphQL query layer + MCP server — zero decorators, zero model changes.
from fastapi import FastAPI
from fastapi_gql_mcp import RouterMCP
app = FastAPI()
# ... your existing routes ...
mcp = RouterMCP(app, name="my-app")
mcp.run() # HTTP MCP server with get_schema + graphql_query tools
Contents — Why · How it works · Installation · Usage · Authentication · Observability · Hardening · Demo · Development · Status
Existing FastAPI→MCP bridges map one tool per endpoint: dozens of tools, no composition, whole-payload responses. fastapi-gql-mcp instead derives a GraphQL schema from your routes (Apollo's "GraphQL as the MCP contract" pattern), so agents get:
{ id name }, not the whole payloaddescription= metadata travel into
the schema the agent readsDepends,
middleware and headers apply; pass credentials via per-caller header
passthrough| Project | Tool count | Field selection | Composition | Setup |
|---|---|---|---|---|
| fastapi-mcp (Tadata) | one per endpoint | ✗ | ✗ | none |
| FastMCP.from_openapi | one per endpoint | ✗ | ✗ | none |
| fastapi-gql-mcp | 2-6, constant | ✓ | ✓ | none |
Full head-to-head — context-growth curves, latency, auth models, selection guidance, all measured on one shared app: comparison/.
comparison/bench/results.json)One app (bench/shared_app.py, a notes CRUD API), wired into both bridges,
driven from two venvs (they can't share one — fastmcp 4 needs mcp>=2,
fastapi-mcp 0.4.0 breaks on mcp 2.x):
| Measurement | fastapi-mcp | fastapi-gql-mcp |
|---|---|---|
| tool catalog, 100 routes | ~10,950 tok (grows linearly) | ~2,480 tok simple / ~1,480 tok progressive (constant) |
| composed task (notes+stats) | 2 tool calls = 2 agent turns | 1 graphql_query |
| same list response | 2,875 B whole payload | 610 B with field projection |
| single trivial call (p50, same client stack, 3 runs) | 0.93 ms | 1.25 ms — GraphQL layer costs; agent turns dominate, not ms |
Honest counterexamples included: below ~5 endpoints the one-tool-per-endpoint
catalog is actually smaller (685 vs 893 tok), and per-call latency favors
them — the GraphQL route pays off as the API grows. Every number is
reproducible (comparison/README.md → Reproduce); environment, versions and
run counts are recorded in results.json.
No decorators, no model changes — everything is derived from the app you already have, in three steps:
RouterScanner reads app.routes: verb, path, params,
response_model, tags, docstrings.Depends, middleware and auth behave exactly as over
HTTP. Sibling fields resolve concurrently; a failing route nulls only its
own field.One route, end to end:
# your code — unchanged
@app.get("/products", response_model=list[ProductOut], tags=["shop:catalog"])
async def list_products(filters: Annotated[ProductFilter, Query()]) -> list[ProductOut]:
"""Browse the product catalog."""
# the schema the agent discovers (excerpt)
type Query { shop: ShopQuery! }
type ShopQuery { catalog: ShopCatalogQuery! }
type ShopCatalogQuery {
list_products(category: String, in_stock: Boolean, limit: Int = 10): [ProductOut!]
}
# what the agent asks — field-level selection, routes combined in one query
{
shop { catalog { list_products(in_stock: true) { name } } }
analytics { shop_stats { revenue_cents } }
}
The shape follows two rules:
tags=["shop:catalog"] answers at
{ shop { catalog { … } } }. Untagged routes join the domain of their first
path segment, so every field has a group.async def list_products becomes
list_products, your own vocabulary with no URL reconstruction. Two routes
sharing a function name fail fast with DuplicateFieldError.Rules worth knowing:
allow_mutation=True to expose writes;
mutation_include=[...] globs to whitelist specific write routes);
graphql_query also refuses mutation documents.JSON — endpoints and fields annotated
dict, dict[K, V] or Any bridge as the JSON scalar instead of being
skipped (both directions: a JSON argument lands as the raw request body).
Routes with NO annotation and no response_model, raw Response returns,
hidden routes and required header/cookie params are still skipped with a
warning — handler.skips lists them programmatically, so CI can assert
nothing fell out of the schema unnoticed.include/exclude fnmatch globs scope which routes enter the schema.tags=["billing:invoice"]); large apps
switch to progressive disclosure (below).Annotated[FilterModel, Query()] flattens into individual query
arguments (FastAPI Query Parameter Models).Field(description=...) → field descriptions, endpoint
docstrings (or summary=) → field descriptions,
Query()/Path()/Body(description=...) → argument descriptions. They surface
in GraphiQL hover, introspection and every MCP discovery tool.Requires Python >= 3.10.
uv add fastapi-gql-mcp # core: GraphQL handler
uv add 'fastapi-gql-mcp[mcp]' # + MCP server (fastmcp)
run() serves streamable HTTP (the only transport — the wrapped app is a
service, and per-caller credential passthrough needs an HTTP request
context). Use mount_to(app, "/mcp") to serve MCP on the app's own port.
mcp = RouterMCP(
app,
name="my-app",
allow_mutation=False,
include=["/api/*"],
# The caller's own Authorization header travels to the routes by default;
# an empty list disables forwarding entirely.
# passthrough_headers=["authorization"],
)
mcp.run() # streamable HTTP, 127.0.0.1:8000 — mcp.run(host="0.0.0.0", port=9000)
Above progressive_threshold routes (default 25, mode="auto"), the toolset
switches to a 4-layer walkthrough of the tag tree:
list_domains ──▶ list_queries("billing:invoice") ──▶ get_query_schema("billing:invoice") ──▶ graphql_query
(list_mutations with allow_mutation=True)
Each domain SDL fragment re-wraps the real group types along the path, so it
shows exactly the grouped query the agent must write — nothing more, and
with every description attached. Discovery is scoped; execution is not:
graphql_query always runs against the full schema, so fields from different
domains combine freely. Force either mode with mode="simple" | "progressive".
mcp.mount_to(app, "/mcp") # streamable HTTP at /mcp/
mcp.handler.mount_graphql(app) # GraphiQL at /graphiql + POST /graphql
from fastapi_gql_mcp import RouterGraphQLHandler
handler = RouterGraphQLHandler(app)
print(handler.get_sdl())
result = await handler.execute(
"query($id: Int!) { iam { get_user(user_id: $id) { name } } }",
variables={"id": 1},
)
Route calls travel through the real ASGI app in-process, so Depends,
middleware and security schemes behave exactly as over HTTP. Credentials have
a single source: the caller — the FastAPI security schemes are the only
verifiers, and this bridge never holds or manages tokens of its own.
passthrough_headers (default
("authorization",)) forwards them to the routes — queries run as the
caller, exactly as they would over HTTP. An explicitly empty list disables
forwarding; headers are matched case-insensitively and only whitelisted
names ever reach a route (no smuggling x-internal-token past the bridge).HTTP_401 in query results; nothing falls back to a server-side identity.handler.execute(..., headers={...}) directly).auth=GitHubProvider(client_id=..., client_secret=..., base_url=...) — and
the MCP endpoint speaks OAuth 2.1: 401 discovery, dynamic client
registration, PKCE, a consent page, and its own reference tokens verifying
every call. Claude Code opens a browser, the user logs in, and the agent's
queries run as that user. mount_to(app, "/mcp", auth_at_root=True) hosts
the OAuth routes at the app root (for reusing an IdP app whose registered
callback lives there). The bridge itself still verifies nothing. Full
wired flow: examples/notes_oauth.Expose the MCP endpoint only behind an entrance you control (network, or a
FastAPI Depends on the mounted route) — the bridge authenticates no one
itself, and combine with allow_mutation=False / mutation_include to keep
writes out of reach.
Install an OpenTelemetry SDK next to your app — that's the whole setup. The spans are emitted natively from both ends, and the bridge stitches them into one waterfall:
tools/call graphql_query);graphql.execute (the GraphQL orchestration layer) and
injects W3C traceparent into every in-process route call — independent
of passthrough_headers, a no-op without an SDK (opentelemetry-api
only, non-recording by default);GET /things plus
fastapi.dependencies/endpoint/serialization) and extracts the injected
context — so route spans nest under graphql.execute, one trace per
query.Route-call timeouts and concurrency queue waits surface as span events
(route.timeout, route.queue) on graphql.execute. A runnable proof
(plus the Jaeger walkthrough): examples/otel_smoke.md;
a live wired app: examples/notes_oauth (env-gated
app/observability.py). Metrics (per-URL QPS/p99) are out of scope here —
derive them from spans with an OTel Collector spanmetrics connector.
Four knobs are built in and on by default:
request_timeout (default 30s, None disables) — per-route-call
deadline. The in-process ASGI call bypasses httpx's own timeout machinery,
so enforcement lives in asyncio.wait_for; a timed-out field surfaces as a
TIMEOUT error (http_status 504) while its siblings survive.max_depth (default 10, None disables) — maximum selection-set
nesting per document. Recursive models make depth unbounded and an MCP
caller is an LLM that can emit runaway nesting; overly deep documents are
rejected with a validation-style error before anything executes.max_concurrency (default 16, None disables) — bound on in-flight
route calls across all queries. Sibling fields resolve concurrently, so
one wide query fans out; this protects the wrapped app's upstream from
being hammered by its own bridge (queueing counts against
request_timeout, default 30s).document_cache_size (default 128, 0 disables) — LRU capacity for the
parse + depth-guard + validate front half of execution, keyed by the query
string. Agents repeat documents constantly; a hit skips straight to
execution (measured 1.52ms → 0.69ms on a 2-field query). Execution results
are never cached — per-call credentials run for real every time.All four are parameters of RouterGraphQLHandler and RouterMCP. For
anything policy-shaped, validation_rules= on the handler passes extra
graphql-core validation rules through (they extend the standard set).
For rate limiting and response caps on the MCP face, FastMCP's
middleware suite attaches with zero bridge code — RouterMCP.mcp is the
underlying FastMCP instance:
from fastmcp.server.middleware.rate_limiting import RateLimitingMiddleware
from fastmcp.server.middleware.response_limiting import ResponseLimitingMiddleware
mcp = RouterMCP(app)
mcp.mcp.add_middleware(RateLimitingMiddleware(max_requests_per_second=10))
mcp.mcp.add_middleware(ResponseLimitingMiddleware(max_size=1_000_000))
RateLimitingMiddleware limits per client by default (pass
get_client_id= to customize the key or global_limit=True for a shared
bucket); ResponseLimitingMiddleware truncates oversized tool responses
(default 1 MB, configurable suffix). The POST /graphql face does not go
through fastmcp — attach your own middleware to the host app for that
endpoint.
The demo/ directory runs a small shop app (users / catalog / orders / stats,
auth via x-token: demo-secret) with every feature in play — including full
documentation coverage so all four description chains are inspectable in
GraphiQL:
uv run --extra mcp python -m demo # REST + /mcp/ + /graphiql + /graphql on :8010
uv run --extra mcp python -m demo.mcp_walkthrough # agent's-eye MCP walkthrough, no client needed
python -m demo prints all endpoint URLs and serves the grouped schema;
/now is untyped on purpose so the skip warning is visible at startup.
For the full consumer experience — a real app with GitHub OAuth login,
session cookies, and MCP OAuth (Claude Code's browser login flow) — see
examples/notes_oauth: three interchangeable
credential carriers resolved in one place, the MCP endpoint protected by
an OAuth 2.1 proxy, and a smoke script that walks the protected paths
headlessly. For observability, examples/otel_smoke.md
walks the one-waterfall-per-query proof in Jaeger.
uv sync && uv run pytest # tests
uv run ruff check src tests
uv run mypy src
0.4.0 — see CHANGELOG.md. Ideas welcome: GraphQL subscriptions over SSE routes, response header pass-through, per-domain auth scopes.
Design extracted from nexusx (SQLModel → GraphQL → MCP), rebuilt on graphql-core standard execution.
MIT
Python
100.0%
FastAPI → GraphQL → MCP: your whole API as ONE constant tool set composed queries, field selection, per-caller auth. Zero Decorators
Python
2
65 commits
updated Oct 5, 2026
Turn any FastAPI router into a GraphQL query layer + MCP server — zero decorators, zero model changes.
from fastapi import FastAPI
from fastapi_gql_mcp import RouterMCP
app = FastAPI()
# ... your existing routes ...
mcp = RouterMCP(app, name="my-app")
mcp.run() # HTTP MCP server with get_schema + graphql_query tools
Contents — Why · How it works · Installation · Usage · Authentication · Observability · Hardening · Demo · Development · Status
Existing FastAPI→MCP bridges map one tool per endpoint: dozens of tools, no composition, whole-payload responses. fastapi-gql-mcp instead derives a GraphQL schema from your routes (Apollo's "GraphQL as the MCP contract" pattern), so agents get:
{ id name }, not the whole payloaddescription= metadata travel into
the schema the agent readsDepends,
middleware and headers apply; pass credentials via per-caller header
passthrough| Project | Tool count | Field selection | Composition | Setup |
|---|---|---|---|---|
| fastapi-mcp (Tadata) | one per endpoint | ✗ | ✗ | none |
| FastMCP.from_openapi | one per endpoint | ✗ | ✗ | none |
| fastapi-gql-mcp | 2-6, constant | ✓ | ✓ | none |
Full head-to-head — context-growth curves, latency, auth models, selection guidance, all measured on one shared app: comparison/.
comparison/bench/results.json)One app (bench/shared_app.py, a notes CRUD API), wired into both bridges,
driven from two venvs (they can't share one — fastmcp 4 needs mcp>=2,
fastapi-mcp 0.4.0 breaks on mcp 2.x):
| Measurement | fastapi-mcp | fastapi-gql-mcp |
|---|---|---|
| tool catalog, 100 routes | ~10,950 tok (grows linearly) | ~2,480 tok simple / ~1,480 tok progressive (constant) |
| composed task (notes+stats) | 2 tool calls = 2 agent turns | 1 graphql_query |
| same list response | 2,875 B whole payload | 610 B with field projection |
| single trivial call (p50, same client stack, 3 runs) | 0.93 ms | 1.25 ms — GraphQL layer costs; agent turns dominate, not ms |
Honest counterexamples included: below ~5 endpoints the one-tool-per-endpoint
catalog is actually smaller (685 vs 893 tok), and per-call latency favors
them — the GraphQL route pays off as the API grows. Every number is
reproducible (comparison/README.md → Reproduce); environment, versions and
run counts are recorded in results.json.
No decorators, no model changes — everything is derived from the app you already have, in three steps:
RouterScanner reads app.routes: verb, path, params,
response_model, tags, docstrings.Depends, middleware and auth behave exactly as over
HTTP. Sibling fields resolve concurrently; a failing route nulls only its
own field.One route, end to end:
# your code — unchanged
@app.get("/products", response_model=list[ProductOut], tags=["shop:catalog"])
async def list_products(filters: Annotated[ProductFilter, Query()]) -> list[ProductOut]:
"""Browse the product catalog."""
# the schema the agent discovers (excerpt)
type Query { shop: ShopQuery! }
type ShopQuery { catalog: ShopCatalogQuery! }
type ShopCatalogQuery {
list_products(category: String, in_stock: Boolean, limit: Int = 10): [ProductOut!]
}
# what the agent asks — field-level selection, routes combined in one query
{
shop { catalog { list_products(in_stock: true) { name } } }
analytics { shop_stats { revenue_cents } }
}
The shape follows two rules:
tags=["shop:catalog"] answers at
{ shop { catalog { … } } }. Untagged routes join the domain of their first
path segment, so every field has a group.async def list_products becomes
list_products, your own vocabulary with no URL reconstruction. Two routes
sharing a function name fail fast with DuplicateFieldError.Rules worth knowing:
allow_mutation=True to expose writes;
mutation_include=[...] globs to whitelist specific write routes);
graphql_query also refuses mutation documents.JSON — endpoints and fields annotated
dict, dict[K, V] or Any bridge as the JSON scalar instead of being
skipped (both directions: a JSON argument lands as the raw request body).
Routes with NO annotation and no response_model, raw Response returns,
hidden routes and required header/cookie params are still skipped with a
warning — handler.skips lists them programmatically, so CI can assert
nothing fell out of the schema unnoticed.include/exclude fnmatch globs scope which routes enter the schema.tags=["billing:invoice"]); large apps
switch to progressive disclosure (below).Annotated[FilterModel, Query()] flattens into individual query
arguments (FastAPI Query Parameter Models).Field(description=...) → field descriptions, endpoint
docstrings (or summary=) → field descriptions,
Query()/Path()/Body(description=...) → argument descriptions. They surface
in GraphiQL hover, introspection and every MCP discovery tool.Requires Python >= 3.10.
uv add fastapi-gql-mcp # core: GraphQL handler
uv add 'fastapi-gql-mcp[mcp]' # + MCP server (fastmcp)
run() serves streamable HTTP (the only transport — the wrapped app is a
service, and per-caller credential passthrough needs an HTTP request
context). Use mount_to(app, "/mcp") to serve MCP on the app's own port.
mcp = RouterMCP(
app,
name="my-app",
allow_mutation=False,
include=["/api/*"],
# The caller's own Authorization header travels to the routes by default;
# an empty list disables forwarding entirely.
# passthrough_headers=["authorization"],
)
mcp.run() # streamable HTTP, 127.0.0.1:8000 — mcp.run(host="0.0.0.0", port=9000)
Above progressive_threshold routes (default 25, mode="auto"), the toolset
switches to a 4-layer walkthrough of the tag tree:
list_domains ──▶ list_queries("billing:invoice") ──▶ get_query_schema("billing:invoice") ──▶ graphql_query
(list_mutations with allow_mutation=True)
Each domain SDL fragment re-wraps the real group types along the path, so it
shows exactly the grouped query the agent must write — nothing more, and
with every description attached. Discovery is scoped; execution is not:
graphql_query always runs against the full schema, so fields from different
domains combine freely. Force either mode with mode="simple" | "progressive".
mcp.mount_to(app, "/mcp") # streamable HTTP at /mcp/
mcp.handler.mount_graphql(app) # GraphiQL at /graphiql + POST /graphql
from fastapi_gql_mcp import RouterGraphQLHandler
handler = RouterGraphQLHandler(app)
print(handler.get_sdl())
result = await handler.execute(
"query($id: Int!) { iam { get_user(user_id: $id) { name } } }",
variables={"id": 1},
)
Route calls travel through the real ASGI app in-process, so Depends,
middleware and security schemes behave exactly as over HTTP. Credentials have
a single source: the caller — the FastAPI security schemes are the only
verifiers, and this bridge never holds or manages tokens of its own.
passthrough_headers (default
("authorization",)) forwards them to the routes — queries run as the
caller, exactly as they would over HTTP. An explicitly empty list disables
forwarding; headers are matched case-insensitively and only whitelisted
names ever reach a route (no smuggling x-internal-token past the bridge).HTTP_401 in query results; nothing falls back to a server-side identity.handler.execute(..., headers={...}) directly).auth=GitHubProvider(client_id=..., client_secret=..., base_url=...) — and
the MCP endpoint speaks OAuth 2.1: 401 discovery, dynamic client
registration, PKCE, a consent page, and its own reference tokens verifying
every call. Claude Code opens a browser, the user logs in, and the agent's
queries run as that user. mount_to(app, "/mcp", auth_at_root=True) hosts
the OAuth routes at the app root (for reusing an IdP app whose registered
callback lives there). The bridge itself still verifies nothing. Full
wired flow: examples/notes_oauth.Expose the MCP endpoint only behind an entrance you control (network, or a
FastAPI Depends on the mounted route) — the bridge authenticates no one
itself, and combine with allow_mutation=False / mutation_include to keep
writes out of reach.
Install an OpenTelemetry SDK next to your app — that's the whole setup. The spans are emitted natively from both ends, and the bridge stitches them into one waterfall:
tools/call graphql_query);graphql.execute (the GraphQL orchestration layer) and
injects W3C traceparent into every in-process route call — independent
of passthrough_headers, a no-op without an SDK (opentelemetry-api
only, non-recording by default);GET /things plus
fastapi.dependencies/endpoint/serialization) and extracts the injected
context — so route spans nest under graphql.execute, one trace per
query.Route-call timeouts and concurrency queue waits surface as span events
(route.timeout, route.queue) on graphql.execute. A runnable proof
(plus the Jaeger walkthrough): examples/otel_smoke.md;
a live wired app: examples/notes_oauth (env-gated
app/observability.py). Metrics (per-URL QPS/p99) are out of scope here —
derive them from spans with an OTel Collector spanmetrics connector.
Four knobs are built in and on by default:
request_timeout (default 30s, None disables) — per-route-call
deadline. The in-process ASGI call bypasses httpx's own timeout machinery,
so enforcement lives in asyncio.wait_for; a timed-out field surfaces as a
TIMEOUT error (http_status 504) while its siblings survive.max_depth (default 10, None disables) — maximum selection-set
nesting per document. Recursive models make depth unbounded and an MCP
caller is an LLM that can emit runaway nesting; overly deep documents are
rejected with a validation-style error before anything executes.max_concurrency (default 16, None disables) — bound on in-flight
route calls across all queries. Sibling fields resolve concurrently, so
one wide query fans out; this protects the wrapped app's upstream from
being hammered by its own bridge (queueing counts against
request_timeout, default 30s).document_cache_size (default 128, 0 disables) — LRU capacity for the
parse + depth-guard + validate front half of execution, keyed by the query
string. Agents repeat documents constantly; a hit skips straight to
execution (measured 1.52ms → 0.69ms on a 2-field query). Execution results
are never cached — per-call credentials run for real every time.All four are parameters of RouterGraphQLHandler and RouterMCP. For
anything policy-shaped, validation_rules= on the handler passes extra
graphql-core validation rules through (they extend the standard set).
For rate limiting and response caps on the MCP face, FastMCP's
middleware suite attaches with zero bridge code — RouterMCP.mcp is the
underlying FastMCP instance:
from fastmcp.server.middleware.rate_limiting import RateLimitingMiddleware
from fastmcp.server.middleware.response_limiting import ResponseLimitingMiddleware
mcp = RouterMCP(app)
mcp.mcp.add_middleware(RateLimitingMiddleware(max_requests_per_second=10))
mcp.mcp.add_middleware(ResponseLimitingMiddleware(max_size=1_000_000))
RateLimitingMiddleware limits per client by default (pass
get_client_id= to customize the key or global_limit=True for a shared
bucket); ResponseLimitingMiddleware truncates oversized tool responses
(default 1 MB, configurable suffix). The POST /graphql face does not go
through fastmcp — attach your own middleware to the host app for that
endpoint.
The demo/ directory runs a small shop app (users / catalog / orders / stats,
auth via x-token: demo-secret) with every feature in play — including full
documentation coverage so all four description chains are inspectable in
GraphiQL:
uv run --extra mcp python -m demo # REST + /mcp/ + /graphiql + /graphql on :8010
uv run --extra mcp python -m demo.mcp_walkthrough # agent's-eye MCP walkthrough, no client needed
python -m demo prints all endpoint URLs and serves the grouped schema;
/now is untyped on purpose so the skip warning is visible at startup.
For the full consumer experience — a real app with GitHub OAuth login,
session cookies, and MCP OAuth (Claude Code's browser login flow) — see
examples/notes_oauth: three interchangeable
credential carriers resolved in one place, the MCP endpoint protected by
an OAuth 2.1 proxy, and a smoke script that walks the protected paths
headlessly. For observability, examples/otel_smoke.md
walks the one-waterfall-per-query proof in Jaeger.
uv sync && uv run pytest # tests
uv run ruff check src tests
uv run mypy src
0.4.0 — see CHANGELOG.md. Ideas welcome: GraphQL subscriptions over SSE routes, response header pass-through, per-domain auth scopes.
Design extracted from nexusx (SQLModel → GraphQL → MCP), rebuilt on graphql-core standard execution.
MIT
Python
100.0%