Kentucky-ai/div9-flooring-estimator-gemma4-31b

Model

div9-flooring-estimator-gemma4-31b (v2.1)

2

5 commits

1 linked in READMEs

updated Jul 13, 2026

See the code

README

div9-flooring-estimator-gemma4-31b (v2.1)

A LoRA adapter for Gemma 4 31B fine-tuned to work as a Division-9 (commercial flooring) bid estimator: given a project header and scoped takeoff quantities, it emits priced line items and a bid total as structured JSON, consistent with one real subcontractor's cost structure.

Headline result: 12.3% median absolute percentage error (APE) on predicted bid totals across a 51-project temporal holdout of real commercial bids β€” every holdout project dated after the training window, every target total verified to the penny against the proposal actually submitted. The untuned base model scores 62.8% median APE on the same gauntlet; a 2024-generation 7B tuned on the same data scores 36.1%.

Weights status: not distributed. The adapter was trained on a company's real, current bid pricing, and the weights necessarily encode that cost structure. This card publishes the method, the recipe, and the verified results. There is no public inference endpoint, so the model cannot be queried for its pricing. A hosted demo is available on request β€” if you are working on estimating models or platforms in adjacent trades, contact via kentucky-ai.com.

The open half of this program

The measurement side of this program is open source. OpenTakeoff (Apache-2.0) is the browser takeoff canvas this model's input contract is built around β€” open a plan, trace rooms with one-click detection, and export exactly the scoped-quantities JSON this model prices. No account, no upload, no seat license.

It also ships the open edition of the capture layer, which is where the training data comes from. Every takeoff an estimator saves stores the drawn regions as normalized vector geometry with their condition labels β€” and reconstructing those polygons reproduces the recorded quantities exactly. Finished estimating work is a labeled, verifiable training corpus, produced at zero marginal labeling cost by the paid work itself. We call the method markup-as-label (patent pending). The boundary is deliberate: the engine and the capture schema are open; the models trained on our own archive β€” this one β€” are the proprietary half.

What we think is actually interesting here

Not the fine-tuning β€” the corpus verification. Anyone can LoRA a 31B on 748 examples. The hard part of "train on your historical bids" is that historical bid documents lie: workbooks descend from templates carrying ghost tabs from other jobs, change-order tabs reference different projects, proposals get misfiled into sibling project folders, and the workbook total is frequently not the number that was submitted. Training targets taken naively from bid workbooks are contaminated in ways no loss curve will reveal.

Every training target here passed a dual-document verification pipeline before admission:

  • Cross-document anchoring. A target bid total must reconcile between the bid workbook and the submitted proposal document. Change-order tabs count only when a real change-order document exists and the tab diffs against its predecessor (template CO tabs from unrelated jobs are direct, observed contamination).
  • Arithmetic forensics. unit_cost Γ— qty is recomputed against every line's extension. Systematic in-band ratios (~1.02–1.25) turn out to be per-line waste/tax factors baked into estimating templates β€” recoverable signal, recorded per line, not noise. Off-band ratios are the true anomalies.
  • Grading. Only projects whose documents reconcile end-to-end (grade A) contribute training targets; per-project verification metrics ship with the corpus.

A practical meta-finding: onboarding a second estimator's project archive initially graded at 63% verified β€” and every disagreement traced to a parser bug, not an estimator error (final: 96% verified, the strongest source in the corpus). A new professional's documents are a fuzz test for your verifier.

Model details

Model typeLoRA adapter (rank 8) over a QAT 4-bit base β€” QLoRA-style: frozen quantized weights, trained low-rank deltas
Base modelmlx-community/gemma-4-31B-it-qat-4bit (Gemma 4 31B instruction-tuned, quantization-aware-trained 4-bit; Apache-2.0)
Trainable parameters32.7 M (0.106% of the 31B base) β€” rank 8, scale 20, dropout 0, top 32 of 60 transformer layers
Language / domainEnglish; commercial flooring (CSI Division 9), US market
Training stackmlx-lm on a single MacBook Pro (128 GB unified memory)
TrainedJuly 2026

Task and I/O contract

Fixed system prompt:

You are a Division-9 (flooring) senior estimator. Given a project and its scoped quantities, produce priced line items and a total consistent with the company's cost structure.

Input (user turn): project header (name, general contractor, date) plus the takeoff β€” scoped quantities as a JSON array, without prices.

Output (assistant turn): one JSON object pricing every line and totaling the bid:

{
  "line_items": [
    {"section": "<string|null>", "name": "<string>", "qty": 0.0,
     "unit": "<string|null>", "cost_type": "Labor|Material|Travel|...",
     "unit_cost": 0.0, "ext": 0.0}
  ],
  "bid_total": 0.0
}

Emitting this schema is itself a tuned behavior: the untuned base writes markdown prose estimates and scores 0/51 under the strict JSON parser. The tuned adapter strict-parses on 48/51 holdout projects with project-specific predictions and clean termination.

Training data

748 supervised pairs from verified historical bids by a single flooring subcontractor: a real project's takeoff quantities paired with the line-item pricing from the bid actually submitted, admitted only through the verification pipeline above. The v2.1 increment over v2 (533 β†’ 748) added a second estimator's archive, with his A-grade projects oversampled 3Γ— as the output-format standard.

The data is not distributed β€” it contains real, current commercial pricing and client-identifying information.

Training recipe

Reproducible on any mlx-lm install against your own corpus:

HyperparameterValue
MethodLoRA (mlx_lm lora --train), gradient checkpointing on
LoRA rank / scale / dropout8 / 20.0 / 0.0
Layers adaptedtop 32 of 60
Batch size2 (no gradient accumulation)
Max sequence length6,144 tokens
Learning rate1e-5, constant (Adam)
Iterations800 scheduled; checkpoint 500 promoted (val loss bottomed at 500, turned by 600)
Peak memory~62 GB
Throughput~35 s/iteration after Metal kernel-compile warmup

Practitioner notes: the first validation pass on a 60-layer graph looks like a hang (~57 min) β€” it is one-time Metal kernel compilation per sequence-length bucket; warm passes take ~10 min. An earlier 7B run in this series collapsed past ~1 epoch on the same task; the 31B QAT base tolerated ~1.9 epochs.

Evaluation

Benchmark: 51 real projects, clean temporal holdout β€” all dated after the training window, targets penny-verified under the corpus' strictest anchoring rules. Metric: median APE of predicted bid_total vs. the verified actual (robust to the long tail).

ModelParamsTraining dataMedian APE≀5%≀10%≀25%Parseable
Qwen2.5-7B base (lenient parse)7Bnone401.6%01151
7B tuned (v0)7B533 pairs36.1%481646
Gemma 4 31B base (lenient parse)31Bnone62.8%14951
Gemma 4 31B tuned (v2)31B533 pairs13.6%11193248
Gemma 4 31B tuned (v2.1, this card)31B748 pairs12.3%12193548

Untuned-base rows use a lenient prose parser (neither base emits the JSON schema unprompted); tuned rows are scored strictly. Both base rows were independently regenerated from scratch with a committed, versioned parser and replicated to three significant figures.

Statistical read (paired bootstrap on both-parsed projects):

  • The base-model swap is the load-bearing effect. 31B-tuned vs. 7B-tuned on identical 533 training examples: better on 35/46 paired projects, median-APE reduction 21.5 points, 95% CI [12.0, 32.4].
  • The data increment is directional, not resolved. v2.1 vs. v2: better on 28/48, median-APE reduction 1.2 points, 95% CI [βˆ’5.2, +7.1] β€” crosses zero at n=51. v2.1 is promoted on point estimates and the ≀25% count, not a significant ablation.
  • Single run per configuration; multi-seed matrix planned. All deltas are preliminary findings.

To our knowledge there is no published multi-project, verified-real-bid APE baseline for this task to compare against; the comparators above are this series' own baselines. A specialty-trade, project-level benchmark built on the same verification methodology is in preparation.

Limitations

  • Estimating aid, not an estimator. Half of predictions miss by more than 12.3%; 16/51 holdout projects missed by >25%. Outputs are a starting point for a human professional, not a submittable bid.
  • One company's cost structure. The model reproduces a single subcontractor's pricing behavior in one regional market, 2024–2026. It does not transfer to other trades, regions, or cost environments, and it goes stale as pricing moves.
  • Garbage-in sensitivity. Inputs are assumed complete, correctly quantified takeoffs; the model does not detect missing scope.
  • No arithmetic guarantee. ext and bid_total are generated, not computed β€” recompute both downstream and discard on disagreement.
  • 3/51 holdout generations failed strict parsing; wrap inference with schema validation and retry.

Contact

Built by Michael Edlin β€” applied AI R&D for the construction trades, Kentucky, USA. Open source: OpenTakeoff β€” the free, Apache-2.0 takeoff engine this program builds on.

apple-silicon
construction
cost-estimation
division-9
estimating
flooring
gemma-4
lora
qlora
text-generation

Kentucky-ai/div9-flooring-estimator-gemma4-31b

Model

div9-flooring-estimator-gemma4-31b (v2.1)

2

5 commits

1 linked in READMEs

updated Jul 13, 2026

See the code

README

div9-flooring-estimator-gemma4-31b (v2.1)

A LoRA adapter for Gemma 4 31B fine-tuned to work as a Division-9 (commercial flooring) bid estimator: given a project header and scoped takeoff quantities, it emits priced line items and a bid total as structured JSON, consistent with one real subcontractor's cost structure.

Headline result: 12.3% median absolute percentage error (APE) on predicted bid totals across a 51-project temporal holdout of real commercial bids β€” every holdout project dated after the training window, every target total verified to the penny against the proposal actually submitted. The untuned base model scores 62.8% median APE on the same gauntlet; a 2024-generation 7B tuned on the same data scores 36.1%.

Weights status: not distributed. The adapter was trained on a company's real, current bid pricing, and the weights necessarily encode that cost structure. This card publishes the method, the recipe, and the verified results. There is no public inference endpoint, so the model cannot be queried for its pricing. A hosted demo is available on request β€” if you are working on estimating models or platforms in adjacent trades, contact via kentucky-ai.com.

The open half of this program

The measurement side of this program is open source. OpenTakeoff (Apache-2.0) is the browser takeoff canvas this model's input contract is built around β€” open a plan, trace rooms with one-click detection, and export exactly the scoped-quantities JSON this model prices. No account, no upload, no seat license.

It also ships the open edition of the capture layer, which is where the training data comes from. Every takeoff an estimator saves stores the drawn regions as normalized vector geometry with their condition labels β€” and reconstructing those polygons reproduces the recorded quantities exactly. Finished estimating work is a labeled, verifiable training corpus, produced at zero marginal labeling cost by the paid work itself. We call the method markup-as-label (patent pending). The boundary is deliberate: the engine and the capture schema are open; the models trained on our own archive β€” this one β€” are the proprietary half.

What we think is actually interesting here

Not the fine-tuning β€” the corpus verification. Anyone can LoRA a 31B on 748 examples. The hard part of "train on your historical bids" is that historical bid documents lie: workbooks descend from templates carrying ghost tabs from other jobs, change-order tabs reference different projects, proposals get misfiled into sibling project folders, and the workbook total is frequently not the number that was submitted. Training targets taken naively from bid workbooks are contaminated in ways no loss curve will reveal.

Every training target here passed a dual-document verification pipeline before admission:

  • Cross-document anchoring. A target bid total must reconcile between the bid workbook and the submitted proposal document. Change-order tabs count only when a real change-order document exists and the tab diffs against its predecessor (template CO tabs from unrelated jobs are direct, observed contamination).
  • Arithmetic forensics. unit_cost Γ— qty is recomputed against every line's extension. Systematic in-band ratios (~1.02–1.25) turn out to be per-line waste/tax factors baked into estimating templates β€” recoverable signal, recorded per line, not noise. Off-band ratios are the true anomalies.
  • Grading. Only projects whose documents reconcile end-to-end (grade A) contribute training targets; per-project verification metrics ship with the corpus.

A practical meta-finding: onboarding a second estimator's project archive initially graded at 63% verified β€” and every disagreement traced to a parser bug, not an estimator error (final: 96% verified, the strongest source in the corpus). A new professional's documents are a fuzz test for your verifier.

Model details

Model typeLoRA adapter (rank 8) over a QAT 4-bit base β€” QLoRA-style: frozen quantized weights, trained low-rank deltas
Base modelmlx-community/gemma-4-31B-it-qat-4bit (Gemma 4 31B instruction-tuned, quantization-aware-trained 4-bit; Apache-2.0)
Trainable parameters32.7 M (0.106% of the 31B base) β€” rank 8, scale 20, dropout 0, top 32 of 60 transformer layers
Language / domainEnglish; commercial flooring (CSI Division 9), US market
Training stackmlx-lm on a single MacBook Pro (128 GB unified memory)
TrainedJuly 2026

Task and I/O contract

Fixed system prompt:

You are a Division-9 (flooring) senior estimator. Given a project and its scoped quantities, produce priced line items and a total consistent with the company's cost structure.

Input (user turn): project header (name, general contractor, date) plus the takeoff β€” scoped quantities as a JSON array, without prices.

Output (assistant turn): one JSON object pricing every line and totaling the bid:

{
  "line_items": [
    {"section": "<string|null>", "name": "<string>", "qty": 0.0,
     "unit": "<string|null>", "cost_type": "Labor|Material|Travel|...",
     "unit_cost": 0.0, "ext": 0.0}
  ],
  "bid_total": 0.0
}

Emitting this schema is itself a tuned behavior: the untuned base writes markdown prose estimates and scores 0/51 under the strict JSON parser. The tuned adapter strict-parses on 48/51 holdout projects with project-specific predictions and clean termination.

Training data

748 supervised pairs from verified historical bids by a single flooring subcontractor: a real project's takeoff quantities paired with the line-item pricing from the bid actually submitted, admitted only through the verification pipeline above. The v2.1 increment over v2 (533 β†’ 748) added a second estimator's archive, with his A-grade projects oversampled 3Γ— as the output-format standard.

The data is not distributed β€” it contains real, current commercial pricing and client-identifying information.

Training recipe

Reproducible on any mlx-lm install against your own corpus:

HyperparameterValue
MethodLoRA (mlx_lm lora --train), gradient checkpointing on
LoRA rank / scale / dropout8 / 20.0 / 0.0
Layers adaptedtop 32 of 60
Batch size2 (no gradient accumulation)
Max sequence length6,144 tokens
Learning rate1e-5, constant (Adam)
Iterations800 scheduled; checkpoint 500 promoted (val loss bottomed at 500, turned by 600)
Peak memory~62 GB
Throughput~35 s/iteration after Metal kernel-compile warmup

Practitioner notes: the first validation pass on a 60-layer graph looks like a hang (~57 min) β€” it is one-time Metal kernel compilation per sequence-length bucket; warm passes take ~10 min. An earlier 7B run in this series collapsed past ~1 epoch on the same task; the 31B QAT base tolerated ~1.9 epochs.

Evaluation

Benchmark: 51 real projects, clean temporal holdout β€” all dated after the training window, targets penny-verified under the corpus' strictest anchoring rules. Metric: median APE of predicted bid_total vs. the verified actual (robust to the long tail).

ModelParamsTraining dataMedian APE≀5%≀10%≀25%Parseable
Qwen2.5-7B base (lenient parse)7Bnone401.6%01151
7B tuned (v0)7B533 pairs36.1%481646
Gemma 4 31B base (lenient parse)31Bnone62.8%14951
Gemma 4 31B tuned (v2)31B533 pairs13.6%11193248
Gemma 4 31B tuned (v2.1, this card)31B748 pairs12.3%12193548

Untuned-base rows use a lenient prose parser (neither base emits the JSON schema unprompted); tuned rows are scored strictly. Both base rows were independently regenerated from scratch with a committed, versioned parser and replicated to three significant figures.

Statistical read (paired bootstrap on both-parsed projects):

  • The base-model swap is the load-bearing effect. 31B-tuned vs. 7B-tuned on identical 533 training examples: better on 35/46 paired projects, median-APE reduction 21.5 points, 95% CI [12.0, 32.4].
  • The data increment is directional, not resolved. v2.1 vs. v2: better on 28/48, median-APE reduction 1.2 points, 95% CI [βˆ’5.2, +7.1] β€” crosses zero at n=51. v2.1 is promoted on point estimates and the ≀25% count, not a significant ablation.
  • Single run per configuration; multi-seed matrix planned. All deltas are preliminary findings.

To our knowledge there is no published multi-project, verified-real-bid APE baseline for this task to compare against; the comparators above are this series' own baselines. A specialty-trade, project-level benchmark built on the same verification methodology is in preparation.

Limitations

  • Estimating aid, not an estimator. Half of predictions miss by more than 12.3%; 16/51 holdout projects missed by >25%. Outputs are a starting point for a human professional, not a submittable bid.
  • One company's cost structure. The model reproduces a single subcontractor's pricing behavior in one regional market, 2024–2026. It does not transfer to other trades, regions, or cost environments, and it goes stale as pricing moves.
  • Garbage-in sensitivity. Inputs are assumed complete, correctly quantified takeoffs; the model does not detect missing scope.
  • No arithmetic guarantee. ext and bid_total are generated, not computed β€” recompute both downstream and discard on disagreement.
  • 3/51 holdout generations failed strict parsing; wrap inference with schema validation and retry.

Contact

Built by Michael Edlin β€” applied AI R&D for the construction trades, Kentucky, USA. Open source: OpenTakeoff β€” the free, Apache-2.0 takeoff engine this program builds on.

apple-silicon
construction
cost-estimation
division-9
estimating
flooring
gemma-4
lora
qlora
text-generation