When exploring complex problems, traditional linear dialogue architectures introduce two fundamental bottlenecks:
GGUF Libra addresses these challenges through a cognitive-inspired design:
Transitions conversational context from a linear log to a version-controlled tree:
Replaces superficial text embeddings with structural topological metrics:
data.bin).Step n: Based on [A] -> Derived [B]) directly into the prompt to reinforce deductive momentum..cpp, .py, .go, and .js source files.llama.cpp instance, evicting the slot's physical KV cache directly from GPU VRAM./props (n_past / n_ctx), raising visual alerts when memory usage exceeds 85%.GGUF Libra relies on llama.cpp for local inference. Two concurrent services are required:
8021).8022).Use the following Python script to launch both services in the background:
import subprocess
import os
import sys
def main():
os.system("")
# Working directory (adjust to your local llama.cpp path)
work_dir = r"D:\llama_cpp"
if not os.path.exists(work_dir):
print(f"Error: Directory not found: {work_dir}")
sys.exit(1)
log_file1_path = os.path.join(work_dir, "llama_main_8021.log")
log_file2_path = os.path.join(work_dir, "llama_embedding_8022.log")
# Service 1: Main chat model (Port 8021)
cmd1 = [
"llama-server.exe",
"-m", r"models\gemma-4-E4B-it-Q5_K_M.gguf",
"--mmproj", r"models\gemma-4-e4b-mmproj-F16.gguf",
"-ngl", "99",
"-c", "32768",
"--port", "8021",
"--host", "0.0.0.0",
"--cache-ram", "512",
"--cache-type-k", "q8_0",
"--cache-type-v", "q8_0",
"--keep", "0",
"-fa", "on",
"-np", "1",
"--reasoning", "off",
"--repeat-penalty", "1.15",
"--mlock"
]
# Service 2: Embedding model (Port 8022)
cmd2 = [
"llama-server.exe",
"-m", r"models\embeddinggemma-300M-Q8_0.gguf",
"-ngl", "99",
"-c", "8192",
"--port", "8022",
"--host", "0.0.0.0",
"--embedding",
"--pooling", "mean"
]
print("Starting llama.cpp services...")
print(f"[Service 1] Main model: http://127.0.0.1:8021")
print(f"[Service 2] Embedding : http://127.0.0.1:8022")
print("Loading models into VRAM...")
try:
log1 = open(log_file1_path, "w", encoding="utf-8")
log2 = open(log_file2_path, "w", encoding="utf-8")
process1 = subprocess.Popen(cmd1, cwd=work_dir, stdout=log1, stderr=subprocess.STDOUT)
process2 = subprocess.Popen(cmd2, cwd=work_dir, stdout=log2, stderr=subprocess.STDOUT)
print("Both services are running in the background.")
print("Press Ctrl + C to terminate both services.\n")
process1.wait()
process2.wait()
except KeyboardInterrupt:
print("\nShutting down services...")
if 'process1' in locals(): process1.terminate()
if 'process2' in locals(): process2.terminate()
print("Services shut down successfully.")
finally:
if 'log1' in locals() and not log1.closed: log1.close()
if 'log2' in locals() and not log2.closed: log2.close()
if __name__ == "__main__":
main()
Run the compiled executable gguf-libra.exe (or go run . during development) and navigate to:
http://127.0.0.1:8099
Use the action bar beneath messages to dynamically manage conversation flow:
To index high-value reasoning trees into the persistent memory store:
cd parser_py
uv run main.py
save: Computes the Graph Laplacian eigenvalues of the active conversation, calculates the topological relaxation weight ($\text{Logic } T$), and persists the sorted derivation tree to data.bin.data: Inspects topological weights and metadata across all indexed trees.From the repository root directory, execute:
# 1. Install frontend dependencies and bundle static distribution
npm install
npm run build
# 2. build parser py
cd parser_py
uv sync
cd ..
# 3. Compile the Go backend binary
go build -x
The resulting gguf-libra.exe runs self-contained alongside the generated dist/ directory.
JavaScript
40.1%
Go
37.0%
HTML
9.6%
Python
9.5%
CSS
3.8%
When exploring complex problems, traditional linear dialogue architectures introduce two fundamental bottlenecks:
GGUF Libra addresses these challenges through a cognitive-inspired design:
Transitions conversational context from a linear log to a version-controlled tree:
Replaces superficial text embeddings with structural topological metrics:
data.bin).Step n: Based on [A] -> Derived [B]) directly into the prompt to reinforce deductive momentum..cpp, .py, .go, and .js source files.llama.cpp instance, evicting the slot's physical KV cache directly from GPU VRAM./props (n_past / n_ctx), raising visual alerts when memory usage exceeds 85%.GGUF Libra relies on llama.cpp for local inference. Two concurrent services are required:
8021).8022).Use the following Python script to launch both services in the background:
import subprocess
import os
import sys
def main():
os.system("")
# Working directory (adjust to your local llama.cpp path)
work_dir = r"D:\llama_cpp"
if not os.path.exists(work_dir):
print(f"Error: Directory not found: {work_dir}")
sys.exit(1)
log_file1_path = os.path.join(work_dir, "llama_main_8021.log")
log_file2_path = os.path.join(work_dir, "llama_embedding_8022.log")
# Service 1: Main chat model (Port 8021)
cmd1 = [
"llama-server.exe",
"-m", r"models\gemma-4-E4B-it-Q5_K_M.gguf",
"--mmproj", r"models\gemma-4-e4b-mmproj-F16.gguf",
"-ngl", "99",
"-c", "32768",
"--port", "8021",
"--host", "0.0.0.0",
"--cache-ram", "512",
"--cache-type-k", "q8_0",
"--cache-type-v", "q8_0",
"--keep", "0",
"-fa", "on",
"-np", "1",
"--reasoning", "off",
"--repeat-penalty", "1.15",
"--mlock"
]
# Service 2: Embedding model (Port 8022)
cmd2 = [
"llama-server.exe",
"-m", r"models\embeddinggemma-300M-Q8_0.gguf",
"-ngl", "99",
"-c", "8192",
"--port", "8022",
"--host", "0.0.0.0",
"--embedding",
"--pooling", "mean"
]
print("Starting llama.cpp services...")
print(f"[Service 1] Main model: http://127.0.0.1:8021")
print(f"[Service 2] Embedding : http://127.0.0.1:8022")
print("Loading models into VRAM...")
try:
log1 = open(log_file1_path, "w", encoding="utf-8")
log2 = open(log_file2_path, "w", encoding="utf-8")
process1 = subprocess.Popen(cmd1, cwd=work_dir, stdout=log1, stderr=subprocess.STDOUT)
process2 = subprocess.Popen(cmd2, cwd=work_dir, stdout=log2, stderr=subprocess.STDOUT)
print("Both services are running in the background.")
print("Press Ctrl + C to terminate both services.\n")
process1.wait()
process2.wait()
except KeyboardInterrupt:
print("\nShutting down services...")
if 'process1' in locals(): process1.terminate()
if 'process2' in locals(): process2.terminate()
print("Services shut down successfully.")
finally:
if 'log1' in locals() and not log1.closed: log1.close()
if 'log2' in locals() and not log2.closed: log2.close()
if __name__ == "__main__":
main()
Run the compiled executable gguf-libra.exe (or go run . during development) and navigate to:
http://127.0.0.1:8099
Use the action bar beneath messages to dynamically manage conversation flow:
To index high-value reasoning trees into the persistent memory store:
cd parser_py
uv run main.py
save: Computes the Graph Laplacian eigenvalues of the active conversation, calculates the topological relaxation weight ($\text{Logic } T$), and persists the sorted derivation tree to data.bin.data: Inspects topological weights and metadata across all indexed trees.From the repository root directory, execute:
# 1. Install frontend dependencies and bundle static distribution
npm install
npm run build
# 2. build parser py
cd parser_py
uv sync
cd ..
# 3. Compile the Go backend binary
go build -x
The resulting gguf-libra.exe runs self-contained alongside the generated dist/ directory.
JavaScript
40.1%
Go
37.0%
HTML
9.6%
Python
9.5%
CSS
3.8%