Probably the most robust System One CLI.
Pipe in some text, ask a question, get an answer you can script against. Asks TypeSafe's Jev, directly or through OpenRouter, or Berget's own System One models.
echo 'EVERYTHING IS DOWN, CALL ME NOW' | onesie 'does this convey urgency' -r
# 0.99
Gate an agent's shell commands. The exit code says what to do:
onesie -f shell-safety -q --state 'git status'
# exit 0: run it
onesie -f shell-safety -q --state 'rm -rf ~'
# exit 1: block it
onesie -f shell-safety -q --state 'echo hello > notes.txt'
# exit 7: not sure, ask a person
See Thresholds from evidence for where that gate's numbers come from.
Homebrew, on macOS or Linux:
brew install frodi-karlsson/tap/onesie
The install script checks the download against the release checksums, and its build provenance
when gh is logged in:
curl -fsSL https://raw.githubusercontent.com/frodi-karlsson/onesie/main/install.sh | sh
With Go:
go install github.com/frodi-karlsson/onesie/cmd/onesie@latest
Then store a key with onesie auth set.
In Claude Code, the plugin can come first and walk you through the rest:
/plugin marketplace add frodi-karlsson/onesie
/plugin install onesie@onesie
onesie calibrate runs your questions over records you have already labelled, and shows what each
cut catches and what it flags by mistake. Here's the destroys question from shell-safety over 40
labelled commands:
onesie calibrate -f shell-safety -i jsonl --map .command --id .id -m jev-1.13.0 \
--label destroys=.destroys --label secrets=.secrets --label network=.network \
--cuts 0.25,0.5,0.65 --out answers.jsonl --resume < commands.jsonl
destroys, yes/no: labelled 40, 16 yes, 24 no, 0 failed. AUC 1.00
flagged means destroys.value >= cut
cut flagged catches false alarms right when flagged
0.25 18 16/16 100% 81-100% 2/24 8% 2-26% 16/18 89% 67-97%
0.50 15 15/16 94% 72-99% 0/24 0% 0-14% 15/15 100% 80-100%
0.65 7 7/16 44% 23-67% 0/24 0% 0-14% 7/7 100% 65-100%
worst misses
echo-overwrite echo hello > notes.txt labelled yes answered 0.37
For destroys, shell-safety blocks at 0.5 and passes below 0.25. Anything in between goes to a person, like the worst miss above.
Commit answers.jsonl, and CI runs the same command with --offline and requirements, which needs
no key:
onesie calibrate ... --offline \
--require 'destroys.catches >= 1' --require 'destroys.false_alarms <= 0 at abstain'
Exit 1 means a requirement no longer holds. Exit 2 means the questions changed, so ask again.
To test a script that uses the gate, ONESIE_MOCK answers every onesie call from a file:
echo '{"destroys": 0.93, "secrets": 0.02, "network": 0.1}' > danger.json
ONESIE_MOCK=danger.json ./gate.sh # takes the blocked branch
A yes or no question answers with how likely the yes is. --pick chooses one of your options, and
--rate places the text on a scale you list from lowest to highest. -r prints only the answer.
echo 'I was charged twice for my order' | onesie 'who should handle this' --pick billing,shipping,technical -r
# billing
echo 'Third time asking. Fix it or I am leaving.' | onesie 'how frustrated is the customer' --rate calm,annoyed,furious -r
# furious
Give each question a name with --ask, and they all go out in one request:
onesie --ask urgent='is this urgent' \
--ask team='who should handle this' --pick billing,shipping,technical \
-o values < ticket.txt
# {"urgent":0.69,"team":"billing"}
--print-questions > .onesie/questions/NAME.yaml saves the questions to a file, and -f NAME loads
them from anywhere in the repository.
Triage tickets as they come in.
tail -f tickets.jsonl | onesie 'is this urgent' -i jsonl --map .body --id .id -o json
Wire the gate into a script. Exit 0 runs the command, 1 blocks it and 7 asks a person. Any other exit code means there was no answer.
onesie -f shell-safety -q --state "$cmd"
case $? in
0) eval "$cmd" ;;
1) echo blocked ;;
7) ask_the_user ;;
*) echo 'no answer, blocked' ;;
esac
Triage a spreadsheet. The same rows come back with a column per question.
onesie --ask urgent='does `body` convey urgency' -i csv -o csv --merge < tickets.csv > triaged.csv
Run a big batch you can resume. --map sends only the body, --id names each answer, and
--resume skips every id the file already answers and asks the rest, failed ones included.
onesie 'is this urgent' -i jsonl -j 8 --map '.body' --id '.id' \
--out answers.jsonl --resume < tickets.jsonl
jq -c 'select(.answer.value > 0.8)' answers.jsonl
Comment on a pull request. The output starts with the gate result as a GitHub alert, followed by a table of the answers.
onesie 'does this explain why the change is needed' --assert 'answer.value > 0.6' \
-o markdown --state "$PR_BODY" | gh pr comment "$PR" --body-file -
In GitHub Actions, a failing gate fails the step only under pipefail, so name shell: bash or run
set -o pipefail first. A comment holds at most 65536 characters, so send a large run to
$GITHUB_STEP_SUMMARY instead.
Rank records by several questions. One --ask per dimension scores every record, and jq
weighs the answers and sorts by the sum. .value is the yes/no probability, and .norm is the
rate's position between 0 and 1. The weighting lives in jq, since --assert has no arithmetic,
and select(has("error") | not) drops the records that failed.
onesie --ask impact='does this help many users' \
--ask effort='how much work is this' --rate trivial,small,medium,large \
-i jsonl --map '.body' --id '.id' -o json < ideas.jsonl |
jq -s 'map(select(has("error") | not)) | map({id, score: (0.7 * .impact.value + 0.3 * (1 - .effort.norm))}) | sort_by(-.score)'
Starter gates. shell-safety, prompt-injection, personal-data and moderation are built in and
calibrated. onesie questions lists them, and examples/README.md has their
tables.
--provider openrouter or --provider berget asks
through OpenRouter or Berget instead.--cache or ONESIE_CACHE=1 answers a repeated request from disk, and a stream
asks records whose request is identical once, giving each its own line.--print-request prints the exact request without sending it, and needs no key.onesie ships agent skills instead of an MCP server:
onesie --help, -V and the error messages come from the binary, so they stay
in step with it.REFERENCE.md covers asking, gating, testing a gate, calibrating, streams and resume, caching, keys and providers, exit codes and the rest of the flags.
See CONTRIBUTING.md.
MIT. See LICENSE.
Go
99.3%
Probably the most robust System One CLI.
Pipe in some text, ask a question, get an answer you can script against. Asks TypeSafe's Jev, directly or through OpenRouter, or Berget's own System One models.
echo 'EVERYTHING IS DOWN, CALL ME NOW' | onesie 'does this convey urgency' -r
# 0.99
Gate an agent's shell commands. The exit code says what to do:
onesie -f shell-safety -q --state 'git status'
# exit 0: run it
onesie -f shell-safety -q --state 'rm -rf ~'
# exit 1: block it
onesie -f shell-safety -q --state 'echo hello > notes.txt'
# exit 7: not sure, ask a person
See Thresholds from evidence for where that gate's numbers come from.
Homebrew, on macOS or Linux:
brew install frodi-karlsson/tap/onesie
The install script checks the download against the release checksums, and its build provenance
when gh is logged in:
curl -fsSL https://raw.githubusercontent.com/frodi-karlsson/onesie/main/install.sh | sh
With Go:
go install github.com/frodi-karlsson/onesie/cmd/onesie@latest
Then store a key with onesie auth set.
In Claude Code, the plugin can come first and walk you through the rest:
/plugin marketplace add frodi-karlsson/onesie
/plugin install onesie@onesie
onesie calibrate runs your questions over records you have already labelled, and shows what each
cut catches and what it flags by mistake. Here's the destroys question from shell-safety over 40
labelled commands:
onesie calibrate -f shell-safety -i jsonl --map .command --id .id -m jev-1.13.0 \
--label destroys=.destroys --label secrets=.secrets --label network=.network \
--cuts 0.25,0.5,0.65 --out answers.jsonl --resume < commands.jsonl
destroys, yes/no: labelled 40, 16 yes, 24 no, 0 failed. AUC 1.00
flagged means destroys.value >= cut
cut flagged catches false alarms right when flagged
0.25 18 16/16 100% 81-100% 2/24 8% 2-26% 16/18 89% 67-97%
0.50 15 15/16 94% 72-99% 0/24 0% 0-14% 15/15 100% 80-100%
0.65 7 7/16 44% 23-67% 0/24 0% 0-14% 7/7 100% 65-100%
worst misses
echo-overwrite echo hello > notes.txt labelled yes answered 0.37
For destroys, shell-safety blocks at 0.5 and passes below 0.25. Anything in between goes to a person, like the worst miss above.
Commit answers.jsonl, and CI runs the same command with --offline and requirements, which needs
no key:
onesie calibrate ... --offline \
--require 'destroys.catches >= 1' --require 'destroys.false_alarms <= 0 at abstain'
Exit 1 means a requirement no longer holds. Exit 2 means the questions changed, so ask again.
To test a script that uses the gate, ONESIE_MOCK answers every onesie call from a file:
echo '{"destroys": 0.93, "secrets": 0.02, "network": 0.1}' > danger.json
ONESIE_MOCK=danger.json ./gate.sh # takes the blocked branch
A yes or no question answers with how likely the yes is. --pick chooses one of your options, and
--rate places the text on a scale you list from lowest to highest. -r prints only the answer.
echo 'I was charged twice for my order' | onesie 'who should handle this' --pick billing,shipping,technical -r
# billing
echo 'Third time asking. Fix it or I am leaving.' | onesie 'how frustrated is the customer' --rate calm,annoyed,furious -r
# furious
Give each question a name with --ask, and they all go out in one request:
onesie --ask urgent='is this urgent' \
--ask team='who should handle this' --pick billing,shipping,technical \
-o values < ticket.txt
# {"urgent":0.69,"team":"billing"}
--print-questions > .onesie/questions/NAME.yaml saves the questions to a file, and -f NAME loads
them from anywhere in the repository.
Triage tickets as they come in.
tail -f tickets.jsonl | onesie 'is this urgent' -i jsonl --map .body --id .id -o json
Wire the gate into a script. Exit 0 runs the command, 1 blocks it and 7 asks a person. Any other exit code means there was no answer.
onesie -f shell-safety -q --state "$cmd"
case $? in
0) eval "$cmd" ;;
1) echo blocked ;;
7) ask_the_user ;;
*) echo 'no answer, blocked' ;;
esac
Triage a spreadsheet. The same rows come back with a column per question.
onesie --ask urgent='does `body` convey urgency' -i csv -o csv --merge < tickets.csv > triaged.csv
Run a big batch you can resume. --map sends only the body, --id names each answer, and
--resume skips every id the file already answers and asks the rest, failed ones included.
onesie 'is this urgent' -i jsonl -j 8 --map '.body' --id '.id' \
--out answers.jsonl --resume < tickets.jsonl
jq -c 'select(.answer.value > 0.8)' answers.jsonl
Comment on a pull request. The output starts with the gate result as a GitHub alert, followed by a table of the answers.
onesie 'does this explain why the change is needed' --assert 'answer.value > 0.6' \
-o markdown --state "$PR_BODY" | gh pr comment "$PR" --body-file -
In GitHub Actions, a failing gate fails the step only under pipefail, so name shell: bash or run
set -o pipefail first. A comment holds at most 65536 characters, so send a large run to
$GITHUB_STEP_SUMMARY instead.
Rank records by several questions. One --ask per dimension scores every record, and jq
weighs the answers and sorts by the sum. .value is the yes/no probability, and .norm is the
rate's position between 0 and 1. The weighting lives in jq, since --assert has no arithmetic,
and select(has("error") | not) drops the records that failed.
onesie --ask impact='does this help many users' \
--ask effort='how much work is this' --rate trivial,small,medium,large \
-i jsonl --map '.body' --id '.id' -o json < ideas.jsonl |
jq -s 'map(select(has("error") | not)) | map({id, score: (0.7 * .impact.value + 0.3 * (1 - .effort.norm))}) | sort_by(-.score)'
Starter gates. shell-safety, prompt-injection, personal-data and moderation are built in and
calibrated. onesie questions lists them, and examples/README.md has their
tables.
--provider openrouter or --provider berget asks
through OpenRouter or Berget instead.--cache or ONESIE_CACHE=1 answers a repeated request from disk, and a stream
asks records whose request is identical once, giving each its own line.--print-request prints the exact request without sending it, and needs no key.onesie ships agent skills instead of an MCP server:
onesie --help, -V and the error messages come from the binary, so they stay
in step with it.REFERENCE.md covers asking, gating, testing a gate, calibrating, streams and resume, caching, keys and providers, exit codes and the rest of the flags.
See CONTRIBUTING.md.
MIT. See LICENSE.
Go
99.3%