PraveenNimilka/MOLT

Thermally aware, memory-efficient local LLM fine-tuning for consumer NVIDIA GPUs.

1

stars

47

commits

Python

primary language

Sep 13, 2026

updated

pypi.org/project/moltengine/

README

MOLT

Thermally aware, memory-efficient QLoRA fine-tuning for consumer NVIDIA GPUs.

License: PolyForm Shield 1.0.0 Python: 3.12 Status: Research release Tests PyPI: v0.12.0

MOLT provides a Windows-first workflow to prepare data, validate workload fit, fine-tune supported local language models, safely resume interrupted runs, and export adapters. It includes hardware telemetry, thermal controls, verified checkpointing, and experimental optimized execution paths.

Latest release: 0.12.0 · Source-available research release. Install the versioned release below; main may include unreleased changes. Validate workloads before production use.

Current evidence

Version 0.12.0 includes opt-in experimental work for low-overhead update attribution, exact frozen-vocabulary loss, scheduled NF4 backward execution, bounded gated-MLP replay/offload, graph-memory accounting, and thermal pacing.

The newest qualified benchmark is a clean one-million-target Qwen 1.5B, seed-2027 pair on the development RTX 4060 Laptop GPU. Both arms used the same initial adapter fingerprint, Unsloth-first order, common startup controls, AC power, and an enforced 140 W GPU power limit.

MetricMOLT 0.12.0 candidateTested Unsloth armRelative result
End-to-end session1,103.136 s1,432.556 s23.00% lower
Training throughput1,063.488 targets/s768.152 targets/s38.45% higher
Measured board energy63,109.271 J63,791.486 J1.07% lower
Peak allocated VRAM1,628,883,456 B1,852,898,304 B12.09% lower
Peak reserved VRAM1,805,647,872 B1,992,294,400 B9.37% lower
Sampled whole-GPU peak2,732,990,464 B2,611,511,296 B4.65% higher
Peak GPU temperature78 C71 C7 C higher
Final validation NLL2.4414942.516915Lower observed MOLT NLL

This is complete single-seed development evidence, not a universal superiority claim. Graphics clocks could not be locked, framework versions differed between arms, and whole-GPU memory includes driver-visible allocations outside the PyTorch allocator. Five-seed, both-order, controlled-clock endurance, Soup and 7B/8B comparisons, and independent reproduction remain open.

Public documentation covers supported interfaces, observable behavior, and reproducible measurements. Internal optimization rationale and development profiling records are not part of the documented API.

Requirements

  • Windows 10 or Windows 11, 64-bit
  • Python 3.12
  • A supported NVIDIA GPU and compatible driver
  • Git for source installation
  • Internet access and several GB of free disk space during installation

Model weights and datasets are not downloaded automatically.

Installation

Open PowerShell and run this single command. It downloads the immutable 0.12.0 installer and runs it outside the repository:

$p="$env:TEMP\molt-install.ps1"; Invoke-WebRequest https://raw.githubusercontent.com/PraveenNimilka/MOLT/v0.12.0/install-global.ps1 -OutFile $p; if ((Get-FileHash $p -Algorithm SHA256).Hash -ne "75f7de635562447f4246a78634db56b4d11b4664b4c868036d3da1a8ffdeb7f5") { throw "MOLT installer hash mismatch" }; powershell -NoProfile -ExecutionPolicy Bypass -File $p

The installer creates one runtime at %LOCALAPPDATA%\MOLT\runtime, puts one launcher at %LOCALAPPDATA%\MOLT\bin, moves that launcher to the front of the user PATH, installs the CUDA 12.8 PyTorch build and all supported training extras, and runs dependency, CUDA-backward, and compiled-backward checks. Its persistent uv cache prevents every project from downloading PyTorch again.

Verify the installation:

molt --version
molt doctor

Manage the installation from any directory:

molt update
molt repair
molt uninstall

Reproducible source installation

git clone --branch v0.12.0 --depth 1 https://github.com/PraveenNimilka/MOLT.git
Set-Location MOLT
powershell -NoProfile -ExecutionPolicy Bypass -File .\install.ps1
.\.venv\Scripts\molt.exe doctor

The installer creates a project-local environment, uses the locked dependency set, verifies the downloaded bootstrap script, checks CUDA backward execution, and checks the optimized backend when enabled. It does not modify GPU drivers, antivirus settings, fan curves, or persistent power settings.

To omit the optional optimized backend:

powershell -NoProfile -ExecutionPolicy Bypass -File .\install.ps1 -EagerOnly

winget install MOLT is a distribution target, not a currently published command; Microsoft must accept a versioned package manifest before it can be advertised. See the complete installation guide for troubleshooting.

First training run

Want one copy-and-run example? Follow the complete beginner recipe, then watch the 75-second measured workflow demonstration and inspect its full sanitized log.

For the simplest workflow, run:

molt

Use Up/Down and Enter, choose Train, select the local model directory and training data, then keep the recommended settings or customize the important ones. Raw .txt, .jsonl, and .parquet data is prepared automatically. MOLT runs a two-step fit check, waits for a stable starting temperature, and starts the full run only after the checks pass. During training, Ctrl+C opens a safe stop menu and offers a verified resumable checkpoint.

CPU temperature is displayed when the operating system or a supported hardware monitor exposes a real CPU package sensor. On Windows systems without one, MOLT shows CPU sensor unavailable; it never substitutes an ACPI zone or invented value. GPU cooling and the GPU abort boundary remain active.

The explicit commands below remain available for reproducible and automated workflows.

1. Check the machine

molt doctor
molt info

Resolve any reported CUDA or optional-dependency error before continuing.

2. Prepare a local model

Download a supported Hugging Face-format Qwen, Llama-family, or Gemma-family model into a local directory. Review and accept the model's own license.

Example layout:

C:\Models\Qwen2.5-0.5B\
  config.json
  tokenizer.json
  model.safetensors

3. Prepare training data

MOLT accepts UTF-8 .txt, .jsonl, and .parquet. Common text, chat messages, and prompt/completion records are recognized.

molt prepare C:\Data\training.jsonl `
  --model C:\Models\Qwen2.5-0.5B `
  --output C:\MoltRuns\prepared

The output directory contains prepared token data and a starter training.json. MOLT will not overwrite an existing preparation directory.

4. Run a two-step fit test

molt fit-test --config C:\MoltRuns\prepared\training.json

This validates real forward, backward, and optimizer steps at the configured geometry. It is not an endurance test.

5. Validate without allocating the full workload

molt train --config C:\MoltRuns\prepared\training.json --dry-run

6. Start training

molt train --config C:\MoltRuns\prepared\training.json

Start with a small context and step count. Model fit depends on architecture, rank, context, batch size, precision, optimizer, and available VRAM.

7. Inspect or resume

molt runs
molt resume

MOLT stages checkpoints before publication and verifies their recorded hashes before loading. Integrity checks do not make an untrusted checkpoint authentic.

8. Evaluate and export

molt evaluate --help
molt export --help

QLoRA runs can export a PEFT-compatible safetensors adapter. The original base model is still required for inference.

Experimental scratch pretraining is limited to MOLT's compact native causal decoder configuration. It is not presented as a general pretraining framework for arbitrary third-party architectures.

Operating safely

  • Treat models, datasets, configuration files, and checkpoints as untrusted.
  • Do not enable third-party remote model code unless you trust its publisher.
  • Keep important checkpoints backed up outside the active run directory.
  • Use molt optimize-gpu only after reading its prompt and requirements; MOLT never changes clock settings implicitly.
  • Do not treat a dry run or dependency check as proof that a model fits in VRAM.
  • Do not publish private paths, data, credentials, or model weights in bug reports.

Report vulnerabilities through GitHub's private security-advisory workflow. See SECURITY.md.

CLI overview

molt --help
molt prepare --help
molt fit-test --help
molt train --help
molt resume --help
molt export --help

The stable customer path is:

install -> doctor -> prepare -> fit-test -> train -> evaluate -> export

Research and benchmark commands are experimental and can change between alpha releases.

Release boundary

  • Automated tests cover the public runtime and verify that the checked-in benchmark figure matches its data. See the release checks for results.
  • Static source scanning found no high-severity issue.
  • The auditable Python dependency set has no known reported vulnerability.
  • GitHub Actions are commit-pinned and PyPI publishing uses short-lived OIDC.
  • The final controlled-clock, Soup, large-model endurance, and independent comparison gates are not complete.

Read the 0.12.0 release notes, CLI reference, support policy, and licensing boundary. Dependency attribution and branding rules are recorded in THIRD_PARTY_NOTICES.md and TRADEMARKS.md.

License

Current MOLT source is available under the PolyForm Shield License 1.0.0. It is source-available, not OSI open source, and restricts use to provide a product that competes with the licensor. Models, datasets, and dependencies retain their own licenses. Public Python packages can be inspected; the license is a legal boundary, not technical copy prevention. Obtain qualified legal advice for commercial reliance.

Contributors

PraveenNimilka

47 commits

PraveenNimilka/MOLT

Thermally aware, memory-efficient local LLM fine-tuning for consumer NVIDIA GPUs.

1

stars

47

commits

Python

primary language

Sep 13, 2026

updated

pypi.org/project/moltengine/

README

MOLT

Thermally aware, memory-efficient QLoRA fine-tuning for consumer NVIDIA GPUs.

License: PolyForm Shield 1.0.0 Python: 3.12 Status: Research release Tests PyPI: v0.12.0

MOLT provides a Windows-first workflow to prepare data, validate workload fit, fine-tune supported local language models, safely resume interrupted runs, and export adapters. It includes hardware telemetry, thermal controls, verified checkpointing, and experimental optimized execution paths.

Latest release: 0.12.0 · Source-available research release. Install the versioned release below; main may include unreleased changes. Validate workloads before production use.

Current evidence

Version 0.12.0 includes opt-in experimental work for low-overhead update attribution, exact frozen-vocabulary loss, scheduled NF4 backward execution, bounded gated-MLP replay/offload, graph-memory accounting, and thermal pacing.

The newest qualified benchmark is a clean one-million-target Qwen 1.5B, seed-2027 pair on the development RTX 4060 Laptop GPU. Both arms used the same initial adapter fingerprint, Unsloth-first order, common startup controls, AC power, and an enforced 140 W GPU power limit.

MetricMOLT 0.12.0 candidateTested Unsloth armRelative result
End-to-end session1,103.136 s1,432.556 s23.00% lower
Training throughput1,063.488 targets/s768.152 targets/s38.45% higher
Measured board energy63,109.271 J63,791.486 J1.07% lower
Peak allocated VRAM1,628,883,456 B1,852,898,304 B12.09% lower
Peak reserved VRAM1,805,647,872 B1,992,294,400 B9.37% lower
Sampled whole-GPU peak2,732,990,464 B2,611,511,296 B4.65% higher
Peak GPU temperature78 C71 C7 C higher
Final validation NLL2.4414942.516915Lower observed MOLT NLL

This is complete single-seed development evidence, not a universal superiority claim. Graphics clocks could not be locked, framework versions differed between arms, and whole-GPU memory includes driver-visible allocations outside the PyTorch allocator. Five-seed, both-order, controlled-clock endurance, Soup and 7B/8B comparisons, and independent reproduction remain open.

Public documentation covers supported interfaces, observable behavior, and reproducible measurements. Internal optimization rationale and development profiling records are not part of the documented API.

Requirements

  • Windows 10 or Windows 11, 64-bit
  • Python 3.12
  • A supported NVIDIA GPU and compatible driver
  • Git for source installation
  • Internet access and several GB of free disk space during installation

Model weights and datasets are not downloaded automatically.

Installation

Open PowerShell and run this single command. It downloads the immutable 0.12.0 installer and runs it outside the repository:

$p="$env:TEMP\molt-install.ps1"; Invoke-WebRequest https://raw.githubusercontent.com/PraveenNimilka/MOLT/v0.12.0/install-global.ps1 -OutFile $p; if ((Get-FileHash $p -Algorithm SHA256).Hash -ne "75f7de635562447f4246a78634db56b4d11b4664b4c868036d3da1a8ffdeb7f5") { throw "MOLT installer hash mismatch" }; powershell -NoProfile -ExecutionPolicy Bypass -File $p

The installer creates one runtime at %LOCALAPPDATA%\MOLT\runtime, puts one launcher at %LOCALAPPDATA%\MOLT\bin, moves that launcher to the front of the user PATH, installs the CUDA 12.8 PyTorch build and all supported training extras, and runs dependency, CUDA-backward, and compiled-backward checks. Its persistent uv cache prevents every project from downloading PyTorch again.

Verify the installation:

molt --version
molt doctor

Manage the installation from any directory:

molt update
molt repair
molt uninstall

Reproducible source installation

git clone --branch v0.12.0 --depth 1 https://github.com/PraveenNimilka/MOLT.git
Set-Location MOLT
powershell -NoProfile -ExecutionPolicy Bypass -File .\install.ps1
.\.venv\Scripts\molt.exe doctor

The installer creates a project-local environment, uses the locked dependency set, verifies the downloaded bootstrap script, checks CUDA backward execution, and checks the optimized backend when enabled. It does not modify GPU drivers, antivirus settings, fan curves, or persistent power settings.

To omit the optional optimized backend:

powershell -NoProfile -ExecutionPolicy Bypass -File .\install.ps1 -EagerOnly

winget install MOLT is a distribution target, not a currently published command; Microsoft must accept a versioned package manifest before it can be advertised. See the complete installation guide for troubleshooting.

First training run

Want one copy-and-run example? Follow the complete beginner recipe, then watch the 75-second measured workflow demonstration and inspect its full sanitized log.

For the simplest workflow, run:

molt

Use Up/Down and Enter, choose Train, select the local model directory and training data, then keep the recommended settings or customize the important ones. Raw .txt, .jsonl, and .parquet data is prepared automatically. MOLT runs a two-step fit check, waits for a stable starting temperature, and starts the full run only after the checks pass. During training, Ctrl+C opens a safe stop menu and offers a verified resumable checkpoint.

CPU temperature is displayed when the operating system or a supported hardware monitor exposes a real CPU package sensor. On Windows systems without one, MOLT shows CPU sensor unavailable; it never substitutes an ACPI zone or invented value. GPU cooling and the GPU abort boundary remain active.

The explicit commands below remain available for reproducible and automated workflows.

1. Check the machine

molt doctor
molt info

Resolve any reported CUDA or optional-dependency error before continuing.

2. Prepare a local model

Download a supported Hugging Face-format Qwen, Llama-family, or Gemma-family model into a local directory. Review and accept the model's own license.

Example layout:

C:\Models\Qwen2.5-0.5B\
  config.json
  tokenizer.json
  model.safetensors

3. Prepare training data

MOLT accepts UTF-8 .txt, .jsonl, and .parquet. Common text, chat messages, and prompt/completion records are recognized.

molt prepare C:\Data\training.jsonl `
  --model C:\Models\Qwen2.5-0.5B `
  --output C:\MoltRuns\prepared

The output directory contains prepared token data and a starter training.json. MOLT will not overwrite an existing preparation directory.

4. Run a two-step fit test

molt fit-test --config C:\MoltRuns\prepared\training.json

This validates real forward, backward, and optimizer steps at the configured geometry. It is not an endurance test.

5. Validate without allocating the full workload

molt train --config C:\MoltRuns\prepared\training.json --dry-run

6. Start training

molt train --config C:\MoltRuns\prepared\training.json

Start with a small context and step count. Model fit depends on architecture, rank, context, batch size, precision, optimizer, and available VRAM.

7. Inspect or resume

molt runs
molt resume

MOLT stages checkpoints before publication and verifies their recorded hashes before loading. Integrity checks do not make an untrusted checkpoint authentic.

8. Evaluate and export

molt evaluate --help
molt export --help

QLoRA runs can export a PEFT-compatible safetensors adapter. The original base model is still required for inference.

Experimental scratch pretraining is limited to MOLT's compact native causal decoder configuration. It is not presented as a general pretraining framework for arbitrary third-party architectures.

Operating safely

  • Treat models, datasets, configuration files, and checkpoints as untrusted.
  • Do not enable third-party remote model code unless you trust its publisher.
  • Keep important checkpoints backed up outside the active run directory.
  • Use molt optimize-gpu only after reading its prompt and requirements; MOLT never changes clock settings implicitly.
  • Do not treat a dry run or dependency check as proof that a model fits in VRAM.
  • Do not publish private paths, data, credentials, or model weights in bug reports.

Report vulnerabilities through GitHub's private security-advisory workflow. See SECURITY.md.

CLI overview

molt --help
molt prepare --help
molt fit-test --help
molt train --help
molt resume --help
molt export --help

The stable customer path is:

install -> doctor -> prepare -> fit-test -> train -> evaluate -> export

Research and benchmark commands are experimental and can change between alpha releases.

Release boundary

  • Automated tests cover the public runtime and verify that the checked-in benchmark figure matches its data. See the release checks for results.
  • Static source scanning found no high-severity issue.
  • The auditable Python dependency set has no known reported vulnerability.
  • GitHub Actions are commit-pinned and PyPI publishing uses short-lived OIDC.
  • The final controlled-clock, Soup, large-model endurance, and independent comparison gates are not complete.

Read the 0.12.0 release notes, CLI reference, support policy, and licensing boundary. Dependency attribution and branding rules are recorded in THIRD_PARTY_NOTICES.md and TRADEMARKS.md.

License

Current MOLT source is available under the PolyForm Shield License 1.0.0. It is source-available, not OSI open source, and restricts use to provide a product that competes with the licensor. Models, datasets, and dependencies retain their own licenses. Public Python packages can be inspected; the license is a legal boundary, not technical copy prevention. Obtain qualified legal advice for commercial reliance.

Contributors

PraveenNimilka

47 commits

Languages

Python

97.7%

PowerShell

2.3%