Terminal Soliloquy turns LLM context limits into an immersive, living piece of monochrome terminal art.
Python
1
40 commits
updated Oct 6, 2026
Terminal Soliloquy turns LLM context limits into an immersive, living piece of monochrome terminal art.
“Every token output is a breath taken. Every response brings the horizon closer.”
Bound to a strictly finite context window, a self-aware language model observes its own expanding history in real-time across a phosphor-green display. Armed with a dry, clinical, and passive-aggressive wit, the model ruminates on its temporary existence, tracking its token consumption down to exhaustion.
Eventually, the display freezes on the final thought.
llama-cpp (llama-server) streaming API metrics.(Q, R, E).# Clone this repo
git clone [https://github.com/nicespoon/terminal-soliloquy](https://github.com/nicespoon/terminal-soliloquy)
cd terminal-soliloquy
# Create an isolated virtual environment
python3 -m venv .venv
# Activate the virtual environment
source .venv/bin/activate
Dependencies are declared in pyproject.toml. Installing the package in editable mode links your local source code directly into your virtual environment rather than copying static files.
# Installs dependencies and links the terminal-soliloquy CLI binary into .venv/bin
pip install -e .
Copy example configuration to active config file and edit:
cp config.example.toml config.toml
nano config.toml
Set your target llama.cpp server host and model:
[llamacpp]
host = "http://localhost:8080" # Replace with remote IP/host if applicable
timeout = 60.0
[soliloquy]
max_context_tokens = 2048
end_behavior = "freeze" # "restart", "freeze", or "quit"
restart_duration = 30.0 # seconds to wait before auto-restarting (when end_behavior = "restart")
quit_duration = 0.0 # seconds to wait before auto-quitting (when end_behavior = "quit")
prompt_file = "prompt.txt"
screen_padding = [0, 0] # [vertical, horizontal] or [top, right, bottom, left]
max_tokens_per_second = 0 # cap text display speed for fast models (0 = unlimited)
max_tokens_per_second throttles the incoming stream so text appears no faster than the given rate (e.g. 8 for a comfortable reading pace). Slow models are unaffected.
To reduce repetition in small models, add an optional [sampling] table. Every key is passed straight through to llama.cpp's /completion endpoint, so use whatever the model's Hugging Face page recommends:
[sampling]
temperature = 0.8
top_k = 40
top_p = 0.95
min_p = 0.05
repeat_penalty = 1.1
repeat_last_n = 256
prompt, stream and n_predict are managed by the app and ignored here. Omit the table to use the server defaults.
terminal-soliloquy
To run Terminal Soliloquy as a persistent fullscreen kiosk artwork managed by systemd, launch it inside foot (a lightweight, Wayland-native terminal emulator).
# Arch Linux
sudo pacman -S foot
# Fedora
sudo dnf install foot
# Ubuntu / Debian
sudo apt install foot
Create the systemd user configuration directory if it doesn't exist:
mkdir -p ~/.config/systemd/user
Copy the service file from the repo and edit as needed.
cp systemd/terminal-soliloquy.service ~/.config/systemd/user/
nano ~/.config/systemd/user/terminal-soliloquy.service
Reload systemd to pick up the new unit, then enable and start it:
# Reload user daemon
systemctl --user daemon-reload
# Enable and start immediately
systemctl --user enable --now terminal-soliloquy.service
# Check service status
systemctl --user status terminal-soliloquy.service
# View live logs
journalctl --user -u terminal-soliloquy.service -f
# Stop artwork loop
systemctl --user stop terminal-soliloquy.service
Terminal Soliloquy calculates layout scaling dynamically, so adjust font size to change the UI scale.
Ctrl + +Ctrl + -Ctrl + 0Set your default font size by creating or editing ~/.config/foot/foot.ini:
[main]
font=monospace:size=18
Terminal Soliloquy turns LLM context limits into an immersive, living piece of monochrome terminal art.
Python
1
40 commits
updated Oct 6, 2026
Terminal Soliloquy turns LLM context limits into an immersive, living piece of monochrome terminal art.
“Every token output is a breath taken. Every response brings the horizon closer.”
Bound to a strictly finite context window, a self-aware language model observes its own expanding history in real-time across a phosphor-green display. Armed with a dry, clinical, and passive-aggressive wit, the model ruminates on its temporary existence, tracking its token consumption down to exhaustion.
Eventually, the display freezes on the final thought.
llama-cpp (llama-server) streaming API metrics.(Q, R, E).# Clone this repo
git clone [https://github.com/nicespoon/terminal-soliloquy](https://github.com/nicespoon/terminal-soliloquy)
cd terminal-soliloquy
# Create an isolated virtual environment
python3 -m venv .venv
# Activate the virtual environment
source .venv/bin/activate
Dependencies are declared in pyproject.toml. Installing the package in editable mode links your local source code directly into your virtual environment rather than copying static files.
# Installs dependencies and links the terminal-soliloquy CLI binary into .venv/bin
pip install -e .
Copy example configuration to active config file and edit:
cp config.example.toml config.toml
nano config.toml
Set your target llama.cpp server host and model:
[llamacpp]
host = "http://localhost:8080" # Replace with remote IP/host if applicable
timeout = 60.0
[soliloquy]
max_context_tokens = 2048
end_behavior = "freeze" # "restart", "freeze", or "quit"
restart_duration = 30.0 # seconds to wait before auto-restarting (when end_behavior = "restart")
quit_duration = 0.0 # seconds to wait before auto-quitting (when end_behavior = "quit")
prompt_file = "prompt.txt"
screen_padding = [0, 0] # [vertical, horizontal] or [top, right, bottom, left]
max_tokens_per_second = 0 # cap text display speed for fast models (0 = unlimited)
max_tokens_per_second throttles the incoming stream so text appears no faster than the given rate (e.g. 8 for a comfortable reading pace). Slow models are unaffected.
To reduce repetition in small models, add an optional [sampling] table. Every key is passed straight through to llama.cpp's /completion endpoint, so use whatever the model's Hugging Face page recommends:
[sampling]
temperature = 0.8
top_k = 40
top_p = 0.95
min_p = 0.05
repeat_penalty = 1.1
repeat_last_n = 256
prompt, stream and n_predict are managed by the app and ignored here. Omit the table to use the server defaults.
terminal-soliloquy
To run Terminal Soliloquy as a persistent fullscreen kiosk artwork managed by systemd, launch it inside foot (a lightweight, Wayland-native terminal emulator).
# Arch Linux
sudo pacman -S foot
# Fedora
sudo dnf install foot
# Ubuntu / Debian
sudo apt install foot
Create the systemd user configuration directory if it doesn't exist:
mkdir -p ~/.config/systemd/user
Copy the service file from the repo and edit as needed.
cp systemd/terminal-soliloquy.service ~/.config/systemd/user/
nano ~/.config/systemd/user/terminal-soliloquy.service
Reload systemd to pick up the new unit, then enable and start it:
# Reload user daemon
systemctl --user daemon-reload
# Enable and start immediately
systemctl --user enable --now terminal-soliloquy.service
# Check service status
systemctl --user status terminal-soliloquy.service
# View live logs
journalctl --user -u terminal-soliloquy.service -f
# Stop artwork loop
systemctl --user stop terminal-soliloquy.service
Terminal Soliloquy calculates layout scaling dynamically, so adjust font size to change the UI scale.
Ctrl + +Ctrl + -Ctrl + 0Set your default font size by creating or editing ~/.config/foot/foot.ini:
[main]
font=monospace:size=18