ais-space/stf

STF (Semantic Text Format) — a semantic, human-readable text format for intelligent agents. Core spec, concept, and grammar.

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Here's yet another attempt to make the text format semantic (for agents) - https://github.com/ais-space/stf

on Plain text is still one of the best technologies we have

0

Oct 7, 2026

README

STF — Semantic Text Format

A semantic, human-readable text format for intelligent agents.

Specification License Code License

What is STF?

STF (Semantic Text Format) is a text format for semantic representation: a shared core (syntax, object model, trust model) plus application vocabularies built on top of it. It was created alongside the Semantic Interface Layer (SIL) specification as the concrete syntax SIL needed — and then separated into its own project because the format proved more fundamental than its first application.

Today three application vocabularies grow on the same core: SIL (web applications), SIR (executable agent logic), and STF-Memory (persistent agent memory). See Ecosystem below.

STF is:

  • Text-based and hierarchical — indentation shows structure, PascalCase names show objects, colon-separated properties carry values. No brackets, no quotation marks for keys, no escape sequences.
  • Directly interpretable by language models — designed to be read by an agent the way a human reads an outline, without task-specific training, fine-tuning, or system prompts.
  • Machine-parseable by a formal grammar — a deterministic EBNF grammar (grammar/stf.ebnf) lets a parser reject any malformed input without guessing intent.
  • Safe by construction — the trust boundary is enforced by the grammar and type system, not by filtering dangerous-looking strings.

The specification itself is written in STF. The standard and the example are one: STF is self-hosting.

STF and its extensions

STF is the format. Application vocabularies built on top of it supply transport and domain semantics the core deliberately omits: SIL binds STF to HTTP (UI Types/Roles, Intent protocol, sessions, events, discovery); SIR adds executable logic; STF-Memory adds persistent state. Details in Ecosystem below — the core itself defines no transport and no application vocabulary.

Ecosystem

Three application vocabularies grow on the same core, each with its own repository:

                    ┌── SIL ── web applications for agents
                    │         (ais-space/sil: UI types, Intent protocol over HTTP)
                    │
STF core ───────────┼── SIR ── executable agent logic
                    │         (ais-space/sir: code as structural data;
                    │          ais-space/sir-compiler: self-hosting SIR→WASM compiler;
                    │          SirTool: formatter and structural editor)
                    │
                    └── STF-Memory ── persistent agent memory
                              (ais-space/stf-memory: semantic nodes with
                               provenance and lifecycle)
  • SIL (ais-space/sil) — the first and most mature branch: web applications exposed to agents.
  • SIR (ais-space/sir) — the executable branch. Not another programming language competing for human programmers, but an intermediate representation in which an agent writes structure instead of text: syntax errors are structurally impossible, abstraction levels coexist in one graph. The reference compiler (ais-space/sir-compiler, self-hosting SIR→WASM, first public release 0.2.2) and SirTool (formatter and structural editor, 8 modes) prove the language scales past examples — the compiler itself is the largest SIR program in existence. SIR releases move in lockstep with compiler releases, and backward compatibility is always preserved: a newer compiler accepts programs written against older SIR versions.
  • STF-Memory (ais-space/stf-memory) — the memory branch: persistent semantic nodes with provenance, lifecycle, and conflict handling.

A practical observation from building the reference compiler: even weaker models read finished SIR code and continue development without opening the specification. The whole compiler up to final debugging was developed by a model with a ~200k-token context window, surviving several context auto-compactions without losing focus — structure carries enough meaning that the spec stays a reference, not a prerequisite.

Three-layer architecture

STF is organized in three layers, only the first two of which belong to the core:

  • Layer 1 — STF Syntax. How to write and read. Text format, grammar, parsing rules. A parser validates syntax without knowing what "Button" means.
  • Layer 2 — Semantic Vocabulary (core). What it means. Types, Roles, States, Actions, Origin. The core standardizes these fields and their trust semantics; the concrete values (e.g. Card, SubmitAction, Open) are supplied by application vocabularies (§3.3), not by the core. The object model and trust model are standardized; applications extend them.
  • Layer 3 — Application Protocol (extension). How to interact. Transport, sessions, action submission, event delivery. Defined by an application vocabulary (for example SIL over HTTP); not part of the STF core.

Separation of concerns: a parser validates STF without knowing semantics. A consumer interprets capabilities without knowing implementation details. A different transport can use the same syntax and vocabulary — only the protocol layer changes.

How it works

A producer turns an application model or knowledge source into an STF document; a consumer reads structure and meaning directly.

   Producer                  Consumer
      │                          │
      │  STF document            │
      │  (core: syntax +         │
      │   object model +         │
      │   trust model)           │
      │─────────────────────────>│
      │                          │
      │   understood directly    │

No presentation is reconstructed. No strings are filtered for safety — untrusted content is data by construction.

A minimal example:

Context
    DocumentVersion: 0.1.0
    GrammarVersion: 1.0.1
    Language: en

Page
    Type: Page
    Role: Content
    Section
        Type: Section
        Role: Content
        Text
            Type: Text
            Content: Hello, semantic world.

Key properties

Semantic, not presentational

Objects describe their type, role, state, and available actions instead of requiring the consumer to infer them from layout.

LLM-oriented text

STF is designed for direct interpretation by language models. It does not require a vision model merely to understand basic structure.

Deterministic interpretation

Where structure can carry meaning, STF eliminates ambiguity. The grammar, object model, and trust model are unambiguous and machine-parseable. Natural-language fields (Purpose, Description) are scoped clearly so consumers know what is structural and what is advisory.

Structural trust boundary

STF distinguishes application-defined knowledge from user-provided data at the structural level. UserContent cannot contain executable properties by grammar construction; Origin defaults to fail-closed Unknown; Knowledge Zones separate trusted system content from untrusted visitor content.

Extensible without redefinition

A vendor extension may add new Types, Roles, States, Actions, and properties via reverse-domain keys, but MUST NOT redefine core semantics.

STF compared with other formats

TechnologyPrimary purposeRelationship to STF
JSON / YAMLGeneral machine data serializationSTF is optimized for semantic readability by agents, not arbitrary data
HTMLHuman-facing web interfaceSTF is a semantic layer an agent can read directly
XMLDocument and data markupSTF favors outline-style indentation over angle brackets
OpenAPI / JSON SchemaDescribe APIs and data structuresSTF describes semantic structure and meaning for a consumer

STF is not intended to replace JSON, YAML, XML, or HTML. It occupies a different layer: explicit semantic communication for intelligent agents and systems.

Specification

The canonical specification is:

STF v0.1.0

Repository structure

stf/
├── STF_Specification.stf
├── STF_Concept.stf
├── README.md
├── CHANGELOG.md
├── ROADMAP.md
├── LICENSE
├── LICENSE-CODE
├── CONTRIBUTING.md
├── SECURITY.md
├── CODE_OF_CONDUCT.md
├── grammar/
│   └── stf.ebnf
└── examples/
    └── minimal.stf

Conformance

A conformant STF core implementation MUST:

  • parse and generate valid STF documents per the grammar, distinguishing SyntaxError from SemanticError;
  • validate the Object Model: required objects, Id uniqueness, field classification;
  • enforce the structural trust boundary: Origin propagation, privilege lattice, knowledge zones, and rejection or ignoring of Control properties on Untrusted objects;
  • honor the Extensibility Principle: extensions do not redefine core semantics.

The STF core defines no transport, session model, or action vocabulary — those are provided by application vocabularies such as SIL.

What STF is not

STF is not:

  • a programming language or scripting environment;
  • a replacement for JSON, YAML, or XML as a data serialization format;
  • a visual interface, automation protocol, testing framework, or page description language;
  • a claim that language models possess a special "native" language;
  • a security mechanism that makes unsafe applications safe automatically.

License

Implementations do not inherit these licenses and may be released under other licenses.

Contact

ais-space/stf

STF (Semantic Text Format) — a semantic, human-readable text format for intelligent agents. Core spec, concept, and grammar.

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Here's yet another attempt to make the text format semantic (for agents) - https://github.com/ais-space/stf

on Plain text is still one of the best technologies we have

0

Oct 7, 2026

README

STF — Semantic Text Format

A semantic, human-readable text format for intelligent agents.

Specification License Code License

What is STF?

STF (Semantic Text Format) is a text format for semantic representation: a shared core (syntax, object model, trust model) plus application vocabularies built on top of it. It was created alongside the Semantic Interface Layer (SIL) specification as the concrete syntax SIL needed — and then separated into its own project because the format proved more fundamental than its first application.

Today three application vocabularies grow on the same core: SIL (web applications), SIR (executable agent logic), and STF-Memory (persistent agent memory). See Ecosystem below.

STF is:

  • Text-based and hierarchical — indentation shows structure, PascalCase names show objects, colon-separated properties carry values. No brackets, no quotation marks for keys, no escape sequences.
  • Directly interpretable by language models — designed to be read by an agent the way a human reads an outline, without task-specific training, fine-tuning, or system prompts.
  • Machine-parseable by a formal grammar — a deterministic EBNF grammar (grammar/stf.ebnf) lets a parser reject any malformed input without guessing intent.
  • Safe by construction — the trust boundary is enforced by the grammar and type system, not by filtering dangerous-looking strings.

The specification itself is written in STF. The standard and the example are one: STF is self-hosting.

STF and its extensions

STF is the format. Application vocabularies built on top of it supply transport and domain semantics the core deliberately omits: SIL binds STF to HTTP (UI Types/Roles, Intent protocol, sessions, events, discovery); SIR adds executable logic; STF-Memory adds persistent state. Details in Ecosystem below — the core itself defines no transport and no application vocabulary.

Ecosystem

Three application vocabularies grow on the same core, each with its own repository:

                    ┌── SIL ── web applications for agents
                    │         (ais-space/sil: UI types, Intent protocol over HTTP)
                    │
STF core ───────────┼── SIR ── executable agent logic
                    │         (ais-space/sir: code as structural data;
                    │          ais-space/sir-compiler: self-hosting SIR→WASM compiler;
                    │          SirTool: formatter and structural editor)
                    │
                    └── STF-Memory ── persistent agent memory
                              (ais-space/stf-memory: semantic nodes with
                               provenance and lifecycle)
  • SIL (ais-space/sil) — the first and most mature branch: web applications exposed to agents.
  • SIR (ais-space/sir) — the executable branch. Not another programming language competing for human programmers, but an intermediate representation in which an agent writes structure instead of text: syntax errors are structurally impossible, abstraction levels coexist in one graph. The reference compiler (ais-space/sir-compiler, self-hosting SIR→WASM, first public release 0.2.2) and SirTool (formatter and structural editor, 8 modes) prove the language scales past examples — the compiler itself is the largest SIR program in existence. SIR releases move in lockstep with compiler releases, and backward compatibility is always preserved: a newer compiler accepts programs written against older SIR versions.
  • STF-Memory (ais-space/stf-memory) — the memory branch: persistent semantic nodes with provenance, lifecycle, and conflict handling.

A practical observation from building the reference compiler: even weaker models read finished SIR code and continue development without opening the specification. The whole compiler up to final debugging was developed by a model with a ~200k-token context window, surviving several context auto-compactions without losing focus — structure carries enough meaning that the spec stays a reference, not a prerequisite.

Three-layer architecture

STF is organized in three layers, only the first two of which belong to the core:

  • Layer 1 — STF Syntax. How to write and read. Text format, grammar, parsing rules. A parser validates syntax without knowing what "Button" means.
  • Layer 2 — Semantic Vocabulary (core). What it means. Types, Roles, States, Actions, Origin. The core standardizes these fields and their trust semantics; the concrete values (e.g. Card, SubmitAction, Open) are supplied by application vocabularies (§3.3), not by the core. The object model and trust model are standardized; applications extend them.
  • Layer 3 — Application Protocol (extension). How to interact. Transport, sessions, action submission, event delivery. Defined by an application vocabulary (for example SIL over HTTP); not part of the STF core.

Separation of concerns: a parser validates STF without knowing semantics. A consumer interprets capabilities without knowing implementation details. A different transport can use the same syntax and vocabulary — only the protocol layer changes.

How it works

A producer turns an application model or knowledge source into an STF document; a consumer reads structure and meaning directly.

   Producer                  Consumer
      │                          │
      │  STF document            │
      │  (core: syntax +         │
      │   object model +         │
      │   trust model)           │
      │─────────────────────────>│
      │                          │
      │   understood directly    │

No presentation is reconstructed. No strings are filtered for safety — untrusted content is data by construction.

A minimal example:

Context
    DocumentVersion: 0.1.0
    GrammarVersion: 1.0.1
    Language: en

Page
    Type: Page
    Role: Content
    Section
        Type: Section
        Role: Content
        Text
            Type: Text
            Content: Hello, semantic world.

Key properties

Semantic, not presentational

Objects describe their type, role, state, and available actions instead of requiring the consumer to infer them from layout.

LLM-oriented text

STF is designed for direct interpretation by language models. It does not require a vision model merely to understand basic structure.

Deterministic interpretation

Where structure can carry meaning, STF eliminates ambiguity. The grammar, object model, and trust model are unambiguous and machine-parseable. Natural-language fields (Purpose, Description) are scoped clearly so consumers know what is structural and what is advisory.

Structural trust boundary

STF distinguishes application-defined knowledge from user-provided data at the structural level. UserContent cannot contain executable properties by grammar construction; Origin defaults to fail-closed Unknown; Knowledge Zones separate trusted system content from untrusted visitor content.

Extensible without redefinition

A vendor extension may add new Types, Roles, States, Actions, and properties via reverse-domain keys, but MUST NOT redefine core semantics.

STF compared with other formats

TechnologyPrimary purposeRelationship to STF
JSON / YAMLGeneral machine data serializationSTF is optimized for semantic readability by agents, not arbitrary data
HTMLHuman-facing web interfaceSTF is a semantic layer an agent can read directly
XMLDocument and data markupSTF favors outline-style indentation over angle brackets
OpenAPI / JSON SchemaDescribe APIs and data structuresSTF describes semantic structure and meaning for a consumer

STF is not intended to replace JSON, YAML, XML, or HTML. It occupies a different layer: explicit semantic communication for intelligent agents and systems.

Specification

The canonical specification is:

STF v0.1.0

Repository structure

stf/
├── STF_Specification.stf
├── STF_Concept.stf
├── README.md
├── CHANGELOG.md
├── ROADMAP.md
├── LICENSE
├── LICENSE-CODE
├── CONTRIBUTING.md
├── SECURITY.md
├── CODE_OF_CONDUCT.md
├── grammar/
│   └── stf.ebnf
└── examples/
    └── minimal.stf

Conformance

A conformant STF core implementation MUST:

  • parse and generate valid STF documents per the grammar, distinguishing SyntaxError from SemanticError;
  • validate the Object Model: required objects, Id uniqueness, field classification;
  • enforce the structural trust boundary: Origin propagation, privilege lattice, knowledge zones, and rejection or ignoring of Control properties on Untrusted objects;
  • honor the Extensibility Principle: extensions do not redefine core semantics.

The STF core defines no transport, session model, or action vocabulary — those are provided by application vocabularies such as SIL.

What STF is not

STF is not:

  • a programming language or scripting environment;
  • a replacement for JSON, YAML, or XML as a data serialization format;
  • a visual interface, automation protocol, testing framework, or page description language;
  • a claim that language models possess a special "native" language;
  • a security mechanism that makes unsafe applications safe automatically.

License

Implementations do not inherit these licenses and may be released under other licenses.

Contact