A domain-specific LLM for cybersecurity β built for ASOC, DFIR and detection-engineering work, and built to run on one machine, inside your own perimeter.
Fine-tuned from Qwen/Qwen3.6-35B-A3B,
a Mixture-of-Experts model with 35B total parameters but only ~3B active per
token. You get large-model domain depth at small-model inference cost: roughly
50 tokens/second on a single NVIDIA DGX Spark, from a 20 GB file.
Released free to the security community by IMPERUM B.V. www.IMPERUM.io, built in partnership with Alican Kiraz. https://huggingface.co/AlicanKiraz0
| File | Size | Notes |
|---|---|---|
Imperum-CybersecurityLLM-v1.0.Q4_K_M.gguf | 20 GB | Recommended. ~22 GB RAM/VRAM. Fastest. |
Imperum-CybersecurityLLM-v1.0.Q8_0.gguf | 34 GB | Near-lossless. ~36 GB RAM/VRAM. |
Everything is inside the .gguf β weights, tokenizer and chat template. No other
files are required.
llama-server -m Imperum-CybersecurityLLM-v1.0.Q4_K_M.gguf \
--port 8099 --ctx-size 16384 --n-gpu-layers 99 \
--jinja --temp 0.3 --reasoning off
Then point any OpenAI client at http://localhost:8099/v1.
ollama create imperum -f Modelfile
ollama run imperum
The bundled Modelfile sets num_ctx 8192 β a safe default for laptops, since
KV cache costs memory on top of the weights. Raise it to 16384 if you have the
headroom.
Drop the .gguf into your models folder and load it. Set temperature to 0.3.
1. It is a reasoning model. The chat template opens a <think> block by
default. If you give it a small max_tokens, the whole budget is consumed by
reasoning and content comes back empty β which looks like a broken
deployment and isn't. Either pass --reasoning off (as above) or allow β₯500
tokens. To get thinking back for a single request, pass
"chat_template_kwargs": {"enable_thinking": true}.
2. Use temperature ~0.3 for security work. The stock default is 1.0, and at 1.0 it will invent plausible-looking CVE numbers and rule syntax. 0.3 is set server-side in the example above; clients can still override per request.
3. It self-identifies as "Qwen". The fine-tune did not change the identity response. Cosmetic, not a misdeploy.
Trained across the full SOC workflow. The examples under each heading were produced by this model on a live deployment (Q8_0, temperature 0.3) β not claimed from the training config.
Turns a raw alert into an analyst-ready assessment: what most likely happened, how confident to be, and what to check before calling it a true positive. Given a Windows event sequence β 4624 logon from an unusual source IP β 4672 SeDebugPrivilege β 7045 service install β 4688 cmd.exe β it reconstructs the intrusion narrative, states a confidence level, and names the single most useful next log source to pull. Handles the classic parent/child cases (winword.exe spawning encoded PowerShell) with the right containment ordering.
Drafts Sigma and YARA content with the correct log source, the correct Sysmon event ID and correct ATT&CK tags attached β a working skeleton for an engineer to review and harden, rather than a blank editor. Also explains why a detection fires and where its false positives will come from.
Maps intrusion chains to MITRE ATT&CK tactics and technique IDs unprompted. Verified correct on T1547.001 (Registry Run Keys), T1053.005 (Scheduled Task), T1204.002 (Malicious File) and T1059.001 (PowerShell). Give it a narrative β phishing ISO β LNK β rundll32 β scheduled task β LSASS dump β SMB lateral movement β staged exfiltration β and it returns the mapped kill chain.
First-hour checklists with the sequencing that actually matters: isolate the host without powering it off, capture volatile memory before disk, preserve event logs before analysis, and establish root cause before cleanup. Covers ransomware, business email compromise and confirmed-intrusion scenarios.
Explains persistence mechanisms and how to hunt them β autostart extensibility points, registry run keys, services, scheduled tasks, Winlogon and IFEO debuggers β along with the tooling (Autoruns, Procmon, Regshot, RegRipper) and the sandbox discipline to do it safely.
Reasons about modern attack paths end to end. From a pod spec with hostPID: true,
privileged: true and a mounted docker socket, it walks container RCE β node root
β kubelet β full cluster compromise, then names the admission policy that blocks
it. Covers Kubernetes, IAM and cloud misconfiguration review.
Identifies vulnerabilities in source, explains the exploitation path, and returns
a fixed version. Correctly flags unsafe pickle.loads() on user-controlled input
as remote code execution and rewrites it safely.
Prioritisation reasoning, remediation guidance, control mapping and the documentation side of security operations β the reporting work that consumes analyst hours without needing analyst judgement.
Intended use: an assistant that accelerates human analysts β drafting, summarising, explaining and triaging. Not an autonomous decision-maker, and not a substitute for review of anything it produces. Detection logic it writes should be reviewed exactly as carefully as human-written logic.
| Property | Value |
|---|---|
| Architecture | qwen3_5_moe (Qwen3.5-MoE) |
| Parameters | 34.66B total, ~3B active per token |
| Layers | 40 |
| Hidden size | 2048 |
| Experts | 256 routed, 8 active per token, plus a shared expert |
| Attention | Hybrid β 3Γ linear attention, then 1Γ full attention, repeating |
| Attention heads | 16 query / 2 key-value, head dim 256 |
| Vocabulary | 248,320 |
| Precision | bfloat16, merged weights (not an adapter) |
| Base | Qwen/Qwen3.6-35B-A3B |
Supervised fine-tune with LoRA r=16, Ξ±=32 across all linear layers, merged into the base weights afterwards. One epoch, learning rate 7e-5 with cosine schedule and 3% warmup, bfloat16, effective batch size 16, 7,750+ optimizer steps, maximum sequence length 5,120. The vision tower was frozen throughout.
Training data covers SOC/SIEM operations, detection engineering, DFIR, malware analysis, threat intelligence, vulnerability management, cloud/Kubernetes/IAM/ AppSec, OT/ICS, GRC and authorized offensive security.
One NVIDIA DGX Spark (GB10 Grace Blackwell, 128 GB unified memory, sm_121), llama.cpp with all layers offloaded to GPU. ~50 tokens/second single stream; ~25 tokens/second per stream with four concurrent requests.
This model is trained for defensive security work: detection, investigation, response and hardening. It also covers authorized offensive security (penetration testing, red teaming) as practised under engagement.
Use it lawfully and only against systems you are authorized to test. As with any LLM, verify its output before acting on it β particularly detection logic, which should be reviewed exactly as carefully as anything written by hand.
Qwen/Qwen3.6-35B-A3B, Apache-2.0).Apache 2.0, inherited from the base model.
@misc{imperum-cybersecurity-llm-v1,
title = {Imperum-CybersecurityLLM-v1.0},
author = {IMPERUM B.V. and Kiraz, Alican},
year = {2026},
note = {Fine-tuned from Qwen/Qwen3.6-35B-A3B},
url = {https://huggingface.co/IMPERUM/Imperum-CybersecurityLLM-v1.0-GGUF}
}
A domain-specific LLM for cybersecurity β built for ASOC, DFIR and detection-engineering work, and built to run on one machine, inside your own perimeter.
Fine-tuned from Qwen/Qwen3.6-35B-A3B,
a Mixture-of-Experts model with 35B total parameters but only ~3B active per
token. You get large-model domain depth at small-model inference cost: roughly
50 tokens/second on a single NVIDIA DGX Spark, from a 20 GB file.
Released free to the security community by IMPERUM B.V. www.IMPERUM.io, built in partnership with Alican Kiraz. https://huggingface.co/AlicanKiraz0
| File | Size | Notes |
|---|---|---|
Imperum-CybersecurityLLM-v1.0.Q4_K_M.gguf | 20 GB | Recommended. ~22 GB RAM/VRAM. Fastest. |
Imperum-CybersecurityLLM-v1.0.Q8_0.gguf | 34 GB | Near-lossless. ~36 GB RAM/VRAM. |
Everything is inside the .gguf β weights, tokenizer and chat template. No other
files are required.
llama-server -m Imperum-CybersecurityLLM-v1.0.Q4_K_M.gguf \
--port 8099 --ctx-size 16384 --n-gpu-layers 99 \
--jinja --temp 0.3 --reasoning off
Then point any OpenAI client at http://localhost:8099/v1.
ollama create imperum -f Modelfile
ollama run imperum
The bundled Modelfile sets num_ctx 8192 β a safe default for laptops, since
KV cache costs memory on top of the weights. Raise it to 16384 if you have the
headroom.
Drop the .gguf into your models folder and load it. Set temperature to 0.3.
1. It is a reasoning model. The chat template opens a <think> block by
default. If you give it a small max_tokens, the whole budget is consumed by
reasoning and content comes back empty β which looks like a broken
deployment and isn't. Either pass --reasoning off (as above) or allow β₯500
tokens. To get thinking back for a single request, pass
"chat_template_kwargs": {"enable_thinking": true}.
2. Use temperature ~0.3 for security work. The stock default is 1.0, and at 1.0 it will invent plausible-looking CVE numbers and rule syntax. 0.3 is set server-side in the example above; clients can still override per request.
3. It self-identifies as "Qwen". The fine-tune did not change the identity response. Cosmetic, not a misdeploy.
Trained across the full SOC workflow. The examples under each heading were produced by this model on a live deployment (Q8_0, temperature 0.3) β not claimed from the training config.
Turns a raw alert into an analyst-ready assessment: what most likely happened, how confident to be, and what to check before calling it a true positive. Given a Windows event sequence β 4624 logon from an unusual source IP β 4672 SeDebugPrivilege β 7045 service install β 4688 cmd.exe β it reconstructs the intrusion narrative, states a confidence level, and names the single most useful next log source to pull. Handles the classic parent/child cases (winword.exe spawning encoded PowerShell) with the right containment ordering.
Drafts Sigma and YARA content with the correct log source, the correct Sysmon event ID and correct ATT&CK tags attached β a working skeleton for an engineer to review and harden, rather than a blank editor. Also explains why a detection fires and where its false positives will come from.
Maps intrusion chains to MITRE ATT&CK tactics and technique IDs unprompted. Verified correct on T1547.001 (Registry Run Keys), T1053.005 (Scheduled Task), T1204.002 (Malicious File) and T1059.001 (PowerShell). Give it a narrative β phishing ISO β LNK β rundll32 β scheduled task β LSASS dump β SMB lateral movement β staged exfiltration β and it returns the mapped kill chain.
First-hour checklists with the sequencing that actually matters: isolate the host without powering it off, capture volatile memory before disk, preserve event logs before analysis, and establish root cause before cleanup. Covers ransomware, business email compromise and confirmed-intrusion scenarios.
Explains persistence mechanisms and how to hunt them β autostart extensibility points, registry run keys, services, scheduled tasks, Winlogon and IFEO debuggers β along with the tooling (Autoruns, Procmon, Regshot, RegRipper) and the sandbox discipline to do it safely.
Reasons about modern attack paths end to end. From a pod spec with hostPID: true,
privileged: true and a mounted docker socket, it walks container RCE β node root
β kubelet β full cluster compromise, then names the admission policy that blocks
it. Covers Kubernetes, IAM and cloud misconfiguration review.
Identifies vulnerabilities in source, explains the exploitation path, and returns
a fixed version. Correctly flags unsafe pickle.loads() on user-controlled input
as remote code execution and rewrites it safely.
Prioritisation reasoning, remediation guidance, control mapping and the documentation side of security operations β the reporting work that consumes analyst hours without needing analyst judgement.
Intended use: an assistant that accelerates human analysts β drafting, summarising, explaining and triaging. Not an autonomous decision-maker, and not a substitute for review of anything it produces. Detection logic it writes should be reviewed exactly as carefully as human-written logic.
| Property | Value |
|---|---|
| Architecture | qwen3_5_moe (Qwen3.5-MoE) |
| Parameters | 34.66B total, ~3B active per token |
| Layers | 40 |
| Hidden size | 2048 |
| Experts | 256 routed, 8 active per token, plus a shared expert |
| Attention | Hybrid β 3Γ linear attention, then 1Γ full attention, repeating |
| Attention heads | 16 query / 2 key-value, head dim 256 |
| Vocabulary | 248,320 |
| Precision | bfloat16, merged weights (not an adapter) |
| Base | Qwen/Qwen3.6-35B-A3B |
Supervised fine-tune with LoRA r=16, Ξ±=32 across all linear layers, merged into the base weights afterwards. One epoch, learning rate 7e-5 with cosine schedule and 3% warmup, bfloat16, effective batch size 16, 7,750+ optimizer steps, maximum sequence length 5,120. The vision tower was frozen throughout.
Training data covers SOC/SIEM operations, detection engineering, DFIR, malware analysis, threat intelligence, vulnerability management, cloud/Kubernetes/IAM/ AppSec, OT/ICS, GRC and authorized offensive security.
One NVIDIA DGX Spark (GB10 Grace Blackwell, 128 GB unified memory, sm_121), llama.cpp with all layers offloaded to GPU. ~50 tokens/second single stream; ~25 tokens/second per stream with four concurrent requests.
This model is trained for defensive security work: detection, investigation, response and hardening. It also covers authorized offensive security (penetration testing, red teaming) as practised under engagement.
Use it lawfully and only against systems you are authorized to test. As with any LLM, verify its output before acting on it β particularly detection logic, which should be reviewed exactly as carefully as anything written by hand.
Qwen/Qwen3.6-35B-A3B, Apache-2.0).Apache 2.0, inherited from the base model.
@misc{imperum-cybersecurity-llm-v1,
title = {Imperum-CybersecurityLLM-v1.0},
author = {IMPERUM B.V. and Kiraz, Alican},
year = {2026},
note = {Fine-tuned from Qwen/Qwen3.6-35B-A3B},
url = {https://huggingface.co/IMPERUM/Imperum-CybersecurityLLM-v1.0-GGUF}
}