div9-flooring-estimator-gemma4-31b (v2.1)
2
5 commits
1 linked in READMEs
updated Jul 13, 2026
A LoRA adapter for Gemma 4 31B fine-tuned to work as a Division-9 (commercial flooring) bid estimator: given a project header and scoped takeoff quantities, it emits priced line items and a bid total as structured JSON, consistent with one real subcontractor's cost structure.
Headline result: 12.3% median absolute percentage error (APE) on predicted bid totals across a 51-project temporal holdout of real commercial bids β every holdout project dated after the training window, every target total verified to the penny against the proposal actually submitted. The untuned base model scores 62.8% median APE on the same gauntlet; a 2024-generation 7B tuned on the same data scores 36.1%.
Weights status: not distributed. The adapter was trained on a company's real, current bid pricing, and the weights necessarily encode that cost structure. This card publishes the method, the recipe, and the verified results. There is no public inference endpoint, so the model cannot be queried for its pricing. A hosted demo is available on request β if you are working on estimating models or platforms in adjacent trades, contact via kentucky-ai.com.
The measurement side of this program is open source. OpenTakeoff (Apache-2.0) is the browser takeoff canvas this model's input contract is built around β open a plan, trace rooms with one-click detection, and export exactly the scoped-quantities JSON this model prices. No account, no upload, no seat license.
It also ships the open edition of the capture layer, which is where the training data comes from. Every takeoff an estimator saves stores the drawn regions as normalized vector geometry with their condition labels β and reconstructing those polygons reproduces the recorded quantities exactly. Finished estimating work is a labeled, verifiable training corpus, produced at zero marginal labeling cost by the paid work itself. We call the method markup-as-label (patent pending). The boundary is deliberate: the engine and the capture schema are open; the models trained on our own archive β this one β are the proprietary half.
Not the fine-tuning β the corpus verification. Anyone can LoRA a 31B on 748 examples. The hard part of "train on your historical bids" is that historical bid documents lie: workbooks descend from templates carrying ghost tabs from other jobs, change-order tabs reference different projects, proposals get misfiled into sibling project folders, and the workbook total is frequently not the number that was submitted. Training targets taken naively from bid workbooks are contaminated in ways no loss curve will reveal.
Every training target here passed a dual-document verification pipeline before admission:
unit_cost Γ qty is recomputed against every
line's extension. Systematic in-band ratios (~1.02β1.25) turn out to be
per-line waste/tax factors baked into estimating templates β recoverable
signal, recorded per line, not noise. Off-band ratios are the true
anomalies.A practical meta-finding: onboarding a second estimator's project archive initially graded at 63% verified β and every disagreement traced to a parser bug, not an estimator error (final: 96% verified, the strongest source in the corpus). A new professional's documents are a fuzz test for your verifier.
| Model type | LoRA adapter (rank 8) over a QAT 4-bit base β QLoRA-style: frozen quantized weights, trained low-rank deltas |
| Base model | mlx-community/gemma-4-31B-it-qat-4bit (Gemma 4 31B instruction-tuned, quantization-aware-trained 4-bit; Apache-2.0) |
| Trainable parameters | 32.7 M (0.106% of the 31B base) β rank 8, scale 20, dropout 0, top 32 of 60 transformer layers |
| Language / domain | English; commercial flooring (CSI Division 9), US market |
| Training stack | mlx-lm on a single MacBook Pro (128 GB unified memory) |
| Trained | July 2026 |
Fixed system prompt:
You are a Division-9 (flooring) senior estimator. Given a project and its scoped quantities, produce priced line items and a total consistent with the company's cost structure.
Input (user turn): project header (name, general contractor, date) plus the takeoff β scoped quantities as a JSON array, without prices.
Output (assistant turn): one JSON object pricing every line and totaling the bid:
{
"line_items": [
{"section": "<string|null>", "name": "<string>", "qty": 0.0,
"unit": "<string|null>", "cost_type": "Labor|Material|Travel|...",
"unit_cost": 0.0, "ext": 0.0}
],
"bid_total": 0.0
}
Emitting this schema is itself a tuned behavior: the untuned base writes markdown prose estimates and scores 0/51 under the strict JSON parser. The tuned adapter strict-parses on 48/51 holdout projects with project-specific predictions and clean termination.
748 supervised pairs from verified historical bids by a single flooring subcontractor: a real project's takeoff quantities paired with the line-item pricing from the bid actually submitted, admitted only through the verification pipeline above. The v2.1 increment over v2 (533 β 748) added a second estimator's archive, with his A-grade projects oversampled 3Γ as the output-format standard.
The data is not distributed β it contains real, current commercial pricing and client-identifying information.
Reproducible on any mlx-lm install against your own corpus:
| Hyperparameter | Value |
|---|---|
| Method | LoRA (mlx_lm lora --train), gradient checkpointing on |
| LoRA rank / scale / dropout | 8 / 20.0 / 0.0 |
| Layers adapted | top 32 of 60 |
| Batch size | 2 (no gradient accumulation) |
| Max sequence length | 6,144 tokens |
| Learning rate | 1e-5, constant (Adam) |
| Iterations | 800 scheduled; checkpoint 500 promoted (val loss bottomed at 500, turned by 600) |
| Peak memory | ~62 GB |
| Throughput | ~35 s/iteration after Metal kernel-compile warmup |
Practitioner notes: the first validation pass on a 60-layer graph looks like a hang (~57 min) β it is one-time Metal kernel compilation per sequence-length bucket; warm passes take ~10 min. An earlier 7B run in this series collapsed past ~1 epoch on the same task; the 31B QAT base tolerated ~1.9 epochs.
Benchmark: 51 real projects, clean temporal holdout β all dated after
the training window, targets penny-verified under the corpus' strictest
anchoring rules. Metric: median APE of predicted bid_total vs. the
verified actual (robust to the long tail).
| Model | Params | Training data | Median APE | β€5% | β€10% | β€25% | Parseable |
|---|---|---|---|---|---|---|---|
| Qwen2.5-7B base (lenient parse) | 7B | none | 401.6% | 0 | 1 | 1 | 51 |
| 7B tuned (v0) | 7B | 533 pairs | 36.1% | 4 | 8 | 16 | 46 |
| Gemma 4 31B base (lenient parse) | 31B | none | 62.8% | 1 | 4 | 9 | 51 |
| Gemma 4 31B tuned (v2) | 31B | 533 pairs | 13.6% | 11 | 19 | 32 | 48 |
| Gemma 4 31B tuned (v2.1, this card) | 31B | 748 pairs | 12.3% | 12 | 19 | 35 | 48 |
Untuned-base rows use a lenient prose parser (neither base emits the JSON schema unprompted); tuned rows are scored strictly. Both base rows were independently regenerated from scratch with a committed, versioned parser and replicated to three significant figures.
Statistical read (paired bootstrap on both-parsed projects):
To our knowledge there is no published multi-project, verified-real-bid APE baseline for this task to compare against; the comparators above are this series' own baselines. A specialty-trade, project-level benchmark built on the same verification methodology is in preparation.
ext and bid_total are generated, not
computed β recompute both downstream and discard on disagreement.Built by Michael Edlin β applied AI R&D for the construction trades, Kentucky, USA. Open source: OpenTakeoff β the free, Apache-2.0 takeoff engine this program builds on.
div9-flooring-estimator-gemma4-31b (v2.1)
2
5 commits
1 linked in READMEs
updated Jul 13, 2026
A LoRA adapter for Gemma 4 31B fine-tuned to work as a Division-9 (commercial flooring) bid estimator: given a project header and scoped takeoff quantities, it emits priced line items and a bid total as structured JSON, consistent with one real subcontractor's cost structure.
Headline result: 12.3% median absolute percentage error (APE) on predicted bid totals across a 51-project temporal holdout of real commercial bids β every holdout project dated after the training window, every target total verified to the penny against the proposal actually submitted. The untuned base model scores 62.8% median APE on the same gauntlet; a 2024-generation 7B tuned on the same data scores 36.1%.
Weights status: not distributed. The adapter was trained on a company's real, current bid pricing, and the weights necessarily encode that cost structure. This card publishes the method, the recipe, and the verified results. There is no public inference endpoint, so the model cannot be queried for its pricing. A hosted demo is available on request β if you are working on estimating models or platforms in adjacent trades, contact via kentucky-ai.com.
The measurement side of this program is open source. OpenTakeoff (Apache-2.0) is the browser takeoff canvas this model's input contract is built around β open a plan, trace rooms with one-click detection, and export exactly the scoped-quantities JSON this model prices. No account, no upload, no seat license.
It also ships the open edition of the capture layer, which is where the training data comes from. Every takeoff an estimator saves stores the drawn regions as normalized vector geometry with their condition labels β and reconstructing those polygons reproduces the recorded quantities exactly. Finished estimating work is a labeled, verifiable training corpus, produced at zero marginal labeling cost by the paid work itself. We call the method markup-as-label (patent pending). The boundary is deliberate: the engine and the capture schema are open; the models trained on our own archive β this one β are the proprietary half.
Not the fine-tuning β the corpus verification. Anyone can LoRA a 31B on 748 examples. The hard part of "train on your historical bids" is that historical bid documents lie: workbooks descend from templates carrying ghost tabs from other jobs, change-order tabs reference different projects, proposals get misfiled into sibling project folders, and the workbook total is frequently not the number that was submitted. Training targets taken naively from bid workbooks are contaminated in ways no loss curve will reveal.
Every training target here passed a dual-document verification pipeline before admission:
unit_cost Γ qty is recomputed against every
line's extension. Systematic in-band ratios (~1.02β1.25) turn out to be
per-line waste/tax factors baked into estimating templates β recoverable
signal, recorded per line, not noise. Off-band ratios are the true
anomalies.A practical meta-finding: onboarding a second estimator's project archive initially graded at 63% verified β and every disagreement traced to a parser bug, not an estimator error (final: 96% verified, the strongest source in the corpus). A new professional's documents are a fuzz test for your verifier.
| Model type | LoRA adapter (rank 8) over a QAT 4-bit base β QLoRA-style: frozen quantized weights, trained low-rank deltas |
| Base model | mlx-community/gemma-4-31B-it-qat-4bit (Gemma 4 31B instruction-tuned, quantization-aware-trained 4-bit; Apache-2.0) |
| Trainable parameters | 32.7 M (0.106% of the 31B base) β rank 8, scale 20, dropout 0, top 32 of 60 transformer layers |
| Language / domain | English; commercial flooring (CSI Division 9), US market |
| Training stack | mlx-lm on a single MacBook Pro (128 GB unified memory) |
| Trained | July 2026 |
Fixed system prompt:
You are a Division-9 (flooring) senior estimator. Given a project and its scoped quantities, produce priced line items and a total consistent with the company's cost structure.
Input (user turn): project header (name, general contractor, date) plus the takeoff β scoped quantities as a JSON array, without prices.
Output (assistant turn): one JSON object pricing every line and totaling the bid:
{
"line_items": [
{"section": "<string|null>", "name": "<string>", "qty": 0.0,
"unit": "<string|null>", "cost_type": "Labor|Material|Travel|...",
"unit_cost": 0.0, "ext": 0.0}
],
"bid_total": 0.0
}
Emitting this schema is itself a tuned behavior: the untuned base writes markdown prose estimates and scores 0/51 under the strict JSON parser. The tuned adapter strict-parses on 48/51 holdout projects with project-specific predictions and clean termination.
748 supervised pairs from verified historical bids by a single flooring subcontractor: a real project's takeoff quantities paired with the line-item pricing from the bid actually submitted, admitted only through the verification pipeline above. The v2.1 increment over v2 (533 β 748) added a second estimator's archive, with his A-grade projects oversampled 3Γ as the output-format standard.
The data is not distributed β it contains real, current commercial pricing and client-identifying information.
Reproducible on any mlx-lm install against your own corpus:
| Hyperparameter | Value |
|---|---|
| Method | LoRA (mlx_lm lora --train), gradient checkpointing on |
| LoRA rank / scale / dropout | 8 / 20.0 / 0.0 |
| Layers adapted | top 32 of 60 |
| Batch size | 2 (no gradient accumulation) |
| Max sequence length | 6,144 tokens |
| Learning rate | 1e-5, constant (Adam) |
| Iterations | 800 scheduled; checkpoint 500 promoted (val loss bottomed at 500, turned by 600) |
| Peak memory | ~62 GB |
| Throughput | ~35 s/iteration after Metal kernel-compile warmup |
Practitioner notes: the first validation pass on a 60-layer graph looks like a hang (~57 min) β it is one-time Metal kernel compilation per sequence-length bucket; warm passes take ~10 min. An earlier 7B run in this series collapsed past ~1 epoch on the same task; the 31B QAT base tolerated ~1.9 epochs.
Benchmark: 51 real projects, clean temporal holdout β all dated after
the training window, targets penny-verified under the corpus' strictest
anchoring rules. Metric: median APE of predicted bid_total vs. the
verified actual (robust to the long tail).
| Model | Params | Training data | Median APE | β€5% | β€10% | β€25% | Parseable |
|---|---|---|---|---|---|---|---|
| Qwen2.5-7B base (lenient parse) | 7B | none | 401.6% | 0 | 1 | 1 | 51 |
| 7B tuned (v0) | 7B | 533 pairs | 36.1% | 4 | 8 | 16 | 46 |
| Gemma 4 31B base (lenient parse) | 31B | none | 62.8% | 1 | 4 | 9 | 51 |
| Gemma 4 31B tuned (v2) | 31B | 533 pairs | 13.6% | 11 | 19 | 32 | 48 |
| Gemma 4 31B tuned (v2.1, this card) | 31B | 748 pairs | 12.3% | 12 | 19 | 35 | 48 |
Untuned-base rows use a lenient prose parser (neither base emits the JSON schema unprompted); tuned rows are scored strictly. Both base rows were independently regenerated from scratch with a committed, versioned parser and replicated to three significant figures.
Statistical read (paired bootstrap on both-parsed projects):
To our knowledge there is no published multi-project, verified-real-bid APE baseline for this task to compare against; the comparators above are this series' own baselines. A specialty-trade, project-level benchmark built on the same verification methodology is in preparation.
ext and bid_total are generated, not
computed β recompute both downstream and discard on disagreement.Built by Michael Edlin β applied AI R&D for the construction trades, Kentucky, USA. Open source: OpenTakeoff β the free, Apache-2.0 takeoff engine this program builds on.