Note: This README is available in English first, followed by the Polish version below.
LLMClient is an advanced, cross-platform AI client built with .NET MAUI that enables seamless interaction with leading language models (LLMs) like GPT and Gemini. The application is designed with security, performance, and flexibility in mind, offering local data storage with encryption and native text tokenization.
tokenizersClone the repository:
git clone https://github.com/DamianTarnowski/LLMClient.git
cd LLMClient
Build the Rust library:
cd TokenizerRust
cargo build --release
This will place the tokenizer_rust.dll file (or equivalent for your system) in the target/release folder. The .NET MAUI project is configured to automatically copy this file during building.
Open and run the .NET MAUI project:
LLMClient.sln file in Visual Studio 2022.Settings screen (gear icon).LLMClient features an advanced memory system that automatically personalizes interactions with AI models:
Automatic Extraction: The system automatically extracts personal information from conversations (name, age, occupation, preferences, etc.) using:
Context in Conversation: Each conversation receives full memory context (maximum 30,000 characters) in the system message, allowing models to provide personalized responses.
Intelligent Management: When memory exceeds the character limit:
Memory Management: An interface is available for viewing, editing, and categorizing saved memories.
The system automatically recognizes and saves:
All memory data is:
The project uses the MVVM pattern and is divided into the following layers:
LLMClient/: Main .NET MAUI project.
Views/: XAML files defining the user interface.ViewModels/: Business logic and state for individual views.Models/: Data models (e.g., Conversation, Message).Services/: Services responsible for key functionalities (e.g., AiService, DatabaseService, EmbeddingService, MemoryContextService, MemoryExtractionService).TokenizerRust/: Rust project containing tokenization logic, exposed as a native library.SafeLocalModelWrapper wraps ILocalModelService, tracks failures, enforces cooldownsRobustLocalModelService, LlamaSharpLocalModelService selectable via EngineSettingsLocalModelDiagnosticService with periodic checks and persisted reportsNetworkAwareDownloadService with auto‑resume and connectivity awarenessMainPageViewModel listens to MessagingCenter events (LocalModelLoaded/LocalModelUnloaded, ModelsChanged), locks cloud picker while local is activeSee the project roadmap for planned enhancements: ROADMAP.md
LLMClient has comprehensive test coverage with 500+ automated tests covering all major functionality.
LLMClient.Tests/
├── Models/ (~163 unit tests)
│ ├── ModelsTests.cs - Conversation, Message, Memory, AiModel
│ ├── EdgeCaseTests.cs - Edge cases and boundary conditions
│ ├── RagTraceTests.cs - RAG tracing and timing
│ ├── ErrorReportTests.cs - Error reporting models
│ ├── EmbeddingModelTests.cs - Embedding model selection
│ ├── DocumentAnalysisTests.cs - Document analysis models
│ ├── EngineSettingsTests.cs - Engine configuration
│ └── ModelSettingsTests.cs - Model settings
├── Services/ (~41 tests)
│ ├── MockServiceTests.cs - Service mocks
│ └── ServiceInterfaceContractTests.cs - Interface contracts
└── Integration/ (~300 tests)
├── MemoryIntegrationTests.cs - Memory CRUD, search, context
├── EmbeddingIntegrationTests.cs - Embedding generation, similarity
├── RagIntegrationTests.cs - RAG document ingestion, retrieval
├── AdvancedRagTests.cs - Hybrid search, reranking
├── ConversationIntegrationTests.cs - Conversation and message handling
├── SearchIntegrationTests.cs - Message search, highlighting
├── LocalModelIntegrationTests.cs - Local model engine, streaming
├── RobustLocalModelTests.cs - Model state, download, memory
├── IngestionPipelineTests.cs - Document queue, processing
├── MemoryContextIntegrationTests.cs - Context generation, summarization
├── AiServiceIntegrationTests.cs - AI generation, summarization
├── EndToEndWorkflowTests.cs - Complete workflows
├── ErrorHandlingIntegrationTests.cs - Error scenarios, retry
├── CacheOptimizationTests.cs - Caching, batching, lazy loading
├── DocumentProcessingTests.cs - Chunking, analysis
├── LocalizationIntegrationTests.cs - Polish/English, formatting
├── SecurityIntegrationTests.cs - API keys, sanitization
├── ExportImportIntegrationTests.cs - Export, backup
├── OnboardingIntegrationTests.cs - First-run experience
├── ModelDownloadIntegrationTests.cs - Download progress, installation
├── ConcurrencyIntegrationTests.cs - Thread safety, rate limiting
├── PerformanceIntegrationTests.cs - Response time, throughput
└── OpenRouterIntegrationTests.cs - Real API integration tests
# Run all unit tests (excluding integration tests with real API)
dotnet test LLMClient.Tests --filter "Category!=Integration"
# Run all tests including integration tests
dotnet test LLMClient.Tests
# Run specific test category
dotnet test LLMClient.Tests --filter "FullyQualifiedName~Memory"
| Area | Tests | Description |
|---|---|---|
| Models | 163 | Data models, enums, properties |
| Memory | 23 | Memory CRUD, search, context building |
| Embeddings | 20 | Generation, similarity, model selection |
| RAG | 28 | Document ingestion, retrieval, hybrid search |
| Conversations | 15 | CRUD, messages, pagination |
| Search | 12 | Highlighting, Polish chars, async |
| Local Models | 28 | State management, download, inference |
| AI Service | 20 | Generation, streaming, summarization |
| Error Handling | 15 | Network, database, retry logic |
| Performance | 13 | Response time, memory, throughput |
| Security | 17 | API keys, sanitization, validation |
| Localization | 12 | Polish/English, date/number formats |
| Export | 9 | JSON, Markdown, backup/restore |
Some tests require an OpenRouter API key stored at ~/.llmclient/openrouter_api_key.txt:
# Create the file with your API key
echo "sk-or-your-api-key" > ~/.llmclient/openrouter_api_key.txt
These tests verify real API behavior including streaming, system prompts, and error handling.
This project is licensed under the MIT license.
LLMClient to zaawansowany, wieloplatformowy klient AI, stworzony w .NET MAUI, który umożliwia płynną interakcję z wiodącymi modelami językowymi (LLM), takimi jak GPT i Gemini. Aplikacja została zaprojektowana z myślą o bezpieczeństwie, wydajności i elastyczności, oferując lokalne przechowywanie danych z szyfrowaniem oraz natywną tokenizację tekstu.
tokenizersSklonuj repozytorium:
git clone https://github.com/DamianTarnowski/LLMClient.git
cd LLMClient
Zbuduj bibliotekę Rust:
cd TokenizerRust
cargo build --release
Spowoduje to umieszczenie pliku tokenizer_rust.dll (lub odpowiednika dla Twojego systemu) w folderze target/release. Projekt .NET MAUI jest skonfigurowany, aby automatycznie kopiować ten plik podczas budowania.
Otwórz i uruchom projekt .NET MAUI:
LLMClient.sln w programie Visual Studio 2022.Ustawienia (ikona koła zębatego).LLMClient posiada zaawansowany system pamięci, który automatycznie personalizuje interakcje z modelami AI:
Automatyczna Ekstraktacja: System automatycznie wydobywa informacje osobiste z konwersacji (imię, wiek, zawód, preferencje, itp.) przy użyciu:
Kontekst w Konwersacji: Każda konwersacja otrzymuje pełny kontekst pamięci (maksymalnie 30,000 znaków) w wiadomości systemowej, pozwalając modelom na spersonalizowane odpowiedzi.
Inteligentne Zarządzanie: Gdy pamięć przekroczy limit znaków:
Zarządzanie Pamięcią: Dostępny jest interfejs do przeglądania, edycji i kategoryzowania zapisanych wspomnień.
System automatycznie rozpoznaje i zapisuje:
Wszystkie dane pamięci są:
Projekt wykorzystuje wzorzec MVVM i jest podzielony na następujące warstwy:
LLMClient/: Główny projekt .NET MAUI.
Views/: Pliki XAML definiujące interfejs użytkownika.ViewModels/: Logika biznesowa i stan dla poszczególnych widoków.Models/: Modele danych (np. Conversation, Message).Services/: Serwisy odpowiedzialne za kluczowe funkcjonalności (np. AiService, DatabaseService, EmbeddingService, MemoryContextService, MemoryExtractionService).TokenizerRust/: Projekt Rust zawierający logikę tokenizacji, udostępnioną jako natywna biblioteka.SafeLocalModelWrapper opakowuje ILocalModelService, śledzi błędy i wymusza cooldownRobustLocalModelService, LlamaSharpLocalModelService wybierane przez EngineSettingsLocalModelDiagnosticService z okresowymi checkami i zapisem raportówNetworkAwareDownloadService z auto‑wznawianiem i świadomością łącznościMainPageViewModel subskrybuje zdarzenia MessagingCenter (LocalModelLoaded/LocalModelUnloaded, ModelsChanged), blokuje wybór chmury gdy lokalny jest aktywnyPlanowane ulepszenia znajdziesz tutaj: ROADMAP.md
LLMClient posiada kompleksowe pokrycie testami z ponad 500 automatycznymi testami obejmującymi wszystkie główne funkcjonalności.
LLMClient.Tests/
├── Models/ (~163 testy jednostkowe)
│ ├── ModelsTests.cs - Conversation, Message, Memory, AiModel
│ ├── EdgeCaseTests.cs - Przypadki brzegowe
│ ├── RagTraceTests.cs - Śledzenie RAG i timing
│ ├── ErrorReportTests.cs - Modele raportowania błędów
│ ├── EmbeddingModelTests.cs - Wybór modelu embeddingowego
│ ├── DocumentAnalysisTests.cs - Modele analizy dokumentów
│ ├── EngineSettingsTests.cs - Konfiguracja silnika
│ └── ModelSettingsTests.cs - Ustawienia modelu
├── Services/ (~41 testów)
│ ├── MockServiceTests.cs - Mocki serwisów
│ └── ServiceInterfaceContractTests.cs - Kontrakty interfejsów
└── Integration/ (~300 testów)
├── MemoryIntegrationTests.cs - Pamięć CRUD, wyszukiwanie
├── EmbeddingIntegrationTests.cs - Generowanie embeddingów
├── RagIntegrationTests.cs - Ingestion dokumentów RAG
├── AdvancedRagTests.cs - Hybrid search, reranking
├── ConversationIntegrationTests.cs - Obsługa konwersacji
├── SearchIntegrationTests.cs - Wyszukiwanie, highlighting
├── LocalModelIntegrationTests.cs - Silnik modeli lokalnych
├── RobustLocalModelTests.cs - Stan modelu, pobieranie
├── IngestionPipelineTests.cs - Kolejka dokumentów
├── MemoryContextIntegrationTests.cs - Generowanie kontekstu
├── AiServiceIntegrationTests.cs - Generowanie AI
├── EndToEndWorkflowTests.cs - Kompletne workflow
├── ErrorHandlingIntegrationTests.cs - Obsługa błędów
├── CacheOptimizationTests.cs - Cache, batching
├── DocumentProcessingTests.cs - Chunking, analiza
├── LocalizationIntegrationTests.cs - PL/EN, formatowanie
├── SecurityIntegrationTests.cs - API keys, sanityzacja
├── ExportImportIntegrationTests.cs - Eksport, backup
├── OnboardingIntegrationTests.cs - Pierwsze uruchomienie
├── ModelDownloadIntegrationTests.cs - Pobieranie modeli
├── ConcurrencyIntegrationTests.cs - Bezpieczeństwo wątków
├── PerformanceIntegrationTests.cs - Wydajność
└── OpenRouterIntegrationTests.cs - Testy z prawdziwym API
# Uruchom testy jednostkowe (bez testów API)
dotnet test LLMClient.Tests --filter "Category!=Integration"
# Uruchom wszystkie testy
dotnet test LLMClient.Tests
# Uruchom konkretną kategorię
dotnet test LLMClient.Tests --filter "FullyQualifiedName~Memory"
| Obszar | Testy | Opis |
|---|---|---|
| Modele | 163 | Modele danych, enumy, właściwości |
| Pamięć | 23 | CRUD, wyszukiwanie, budowanie kontekstu |
| Embeddingi | 20 | Generowanie, podobieństwo, wybór modelu |
| RAG | 28 | Ingestion, retrieval, hybrid search |
| Konwersacje | 15 | CRUD, wiadomości, paginacja |
| Wyszukiwanie | 12 | Highlighting, polskie znaki |
| Modele lokalne | 28 | Zarządzanie stanem, pobieranie |
| AI Service | 20 | Generowanie, streaming |
| Błędy | 15 | Sieć, baza danych, retry |
| Wydajność | 13 | Czas odpowiedzi, pamięć |
| Bezpieczeństwo | 17 | API keys, sanityzacja |
| Lokalizacja | 12 | PL/EN, formaty dat/liczb |
| Eksport | 9 | JSON, Markdown, backup |
Niektóre testy wymagają klucza API OpenRouter w ~/.llmclient/openrouter_api_key.txt:
# Utwórz plik z kluczem API
echo "sk-or-your-api-key" > ~/.llmclient/openrouter_api_key.txt
Ten projekt jest udostępniany na licencji MIT.
LLMClient offers two embedding models for semantic search and RAG (Retrieval-Augmented Generation):
| Model | Quality | Speed | Size | Min RAM | Best For |
|---|---|---|---|---|---|
| E5-Large Multilingual | 🎯 95% | 60% | ~2.2GB | 8GB | Desktop, high-quality search |
| Gemma 300M | 🚀 75% | 90% | ~1.2GB | 4GB | Mobile, resource-constrained devices |
The app automatically selects the best model based on device RAM:
Users can override this in Settings → Embedding Model.
Tests performed on Polish language pairs comparing semantic similarity scores:
| Pair | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Uwielbiam kodować" | 0.936 | 0.840 | +11% |
| "Kot śpi" ↔ "Kotek drzemie" | 0.944 | 0.614 | +54% |
| "Auto jedzie" ↔ "Samochód pędzi" | 0.952 | 0.682 | +40% |
| Average | 0.944 | 0.712 | +33% |
| Pair | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Pierogi są pyszne" | 0.829 | 0.441 | -47% |
| "Kot śpi" ↔ "Samolot leci" | 0.833 | 0.440 | -47% |
| "Auto jedzie" ↔ "Książka leży" | 0.788 | 0.393 | -50% |
| Average | 0.817 | 0.425 | -48% |
| Pair | E5-Large | Gemma 300M |
|---|---|---|
| PL: "Lubię programować" ↔ EN: "I like programming" | 0.966 | 0.701 |
| PL: "Kot jest zwierzęciem" ↔ EN: "A cat is an animal" | 0.920 | 0.794 |
| Average | 0.943 | 0.748 |
E5-Large excels at Polish semantic similarity - 33% higher scores for similar sentences, making it ideal for precise semantic search in Polish documents.
Gemma better discriminates unrelated content - 48% lower scores for different sentences means fewer false positives in search results.
E5-Large superior for cross-lingual - Critical for applications mixing Polish and English content.
Trade-off: E5-Large requires 2x more RAM but provides significantly better quality for Polish language semantic search. For mobile devices with <8GB RAM, Gemma offers acceptable quality with much better performance.
Both models use native Rust tokenizers for optimal performance:
How to use the local model
LLMClient oferuje dwa modele embeddingowe do wyszukiwania semantycznego i RAG:
| Model | Jakość | Szybkość | Rozmiar | Min RAM | Najlepszy dla |
|---|---|---|---|---|---|
| E5-Large Multilingual | 🎯 95% | 60% | ~2.2GB | 8GB | Desktop, wysokiej jakości wyszukiwanie |
| Gemma 300M | 🚀 75% | 90% | ~1.2GB | 4GB | Mobile, urządzenia z ograniczeniami |
Aplikacja automatycznie wybiera najlepszy model na podstawie RAM urządzenia:
Użytkownik może to zmienić w Ustawieniach → Model Embeddingowy.
Testy przeprowadzone na polskich parach zdań porównujące wyniki podobieństwa semantycznego:
| Para | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Uwielbiam kodować" | 0.936 | 0.840 | +11% |
| "Kot śpi" ↔ "Kotek drzemie" | 0.944 | 0.614 | +54% |
| "Auto jedzie" ↔ "Samochód pędzi" | 0.952 | 0.682 | +40% |
| Średnia | 0.944 | 0.712 | +33% |
| Para | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Pierogi są pyszne" | 0.829 | 0.441 | -47% |
| "Kot śpi" ↔ "Samolot leci" | 0.833 | 0.440 | -47% |
| "Auto jedzie" ↔ "Książka leży" | 0.788 | 0.393 | -50% |
| Średnia | 0.817 | 0.425 | -48% |
| Para | E5-Large | Gemma 300M |
|---|---|---|
| PL: "Lubię programować" ↔ EN: "I like programming" | 0.966 | 0.701 |
| PL: "Kot jest zwierzęciem" ↔ EN: "A cat is an animal" | 0.920 | 0.794 |
| Średnia | 0.943 | 0.748 |
E5-Large doskonały dla polskiego podobieństwa semantycznego - 33% wyższe wyniki dla podobnych zdań, idealny do precyzyjnego wyszukiwania semantycznego w polskich dokumentach.
Gemma lepiej rozróżnia niepowiązane treści - 48% niższe wyniki dla różnych zdań oznacza mniej fałszywych trafień w wynikach wyszukiwania.
E5-Large przewyższa w cross-lingual - Krytyczne dla aplikacji mieszających polskie i angielskie treści.
Kompromis: E5-Large wymaga 2x więcej RAM, ale zapewnia znacznie lepszą jakość dla polskiego wyszukiwania semantycznego. Dla urządzeń mobilnych z <8GB RAM, Gemma oferuje akceptowalną jakość przy znacznie lepszej wydajności.
Oba modele używają natywnych tokenizerów Rust dla optymalnej wydajności:
Jak korzystać z modelu lokalnego
C#
98.0%
Rust
1.4%
Note: This README is available in English first, followed by the Polish version below.
LLMClient is an advanced, cross-platform AI client built with .NET MAUI that enables seamless interaction with leading language models (LLMs) like GPT and Gemini. The application is designed with security, performance, and flexibility in mind, offering local data storage with encryption and native text tokenization.
tokenizersClone the repository:
git clone https://github.com/DamianTarnowski/LLMClient.git
cd LLMClient
Build the Rust library:
cd TokenizerRust
cargo build --release
This will place the tokenizer_rust.dll file (or equivalent for your system) in the target/release folder. The .NET MAUI project is configured to automatically copy this file during building.
Open and run the .NET MAUI project:
LLMClient.sln file in Visual Studio 2022.Settings screen (gear icon).LLMClient features an advanced memory system that automatically personalizes interactions with AI models:
Automatic Extraction: The system automatically extracts personal information from conversations (name, age, occupation, preferences, etc.) using:
Context in Conversation: Each conversation receives full memory context (maximum 30,000 characters) in the system message, allowing models to provide personalized responses.
Intelligent Management: When memory exceeds the character limit:
Memory Management: An interface is available for viewing, editing, and categorizing saved memories.
The system automatically recognizes and saves:
All memory data is:
The project uses the MVVM pattern and is divided into the following layers:
LLMClient/: Main .NET MAUI project.
Views/: XAML files defining the user interface.ViewModels/: Business logic and state for individual views.Models/: Data models (e.g., Conversation, Message).Services/: Services responsible for key functionalities (e.g., AiService, DatabaseService, EmbeddingService, MemoryContextService, MemoryExtractionService).TokenizerRust/: Rust project containing tokenization logic, exposed as a native library.SafeLocalModelWrapper wraps ILocalModelService, tracks failures, enforces cooldownsRobustLocalModelService, LlamaSharpLocalModelService selectable via EngineSettingsLocalModelDiagnosticService with periodic checks and persisted reportsNetworkAwareDownloadService with auto‑resume and connectivity awarenessMainPageViewModel listens to MessagingCenter events (LocalModelLoaded/LocalModelUnloaded, ModelsChanged), locks cloud picker while local is activeSee the project roadmap for planned enhancements: ROADMAP.md
LLMClient has comprehensive test coverage with 500+ automated tests covering all major functionality.
LLMClient.Tests/
├── Models/ (~163 unit tests)
│ ├── ModelsTests.cs - Conversation, Message, Memory, AiModel
│ ├── EdgeCaseTests.cs - Edge cases and boundary conditions
│ ├── RagTraceTests.cs - RAG tracing and timing
│ ├── ErrorReportTests.cs - Error reporting models
│ ├── EmbeddingModelTests.cs - Embedding model selection
│ ├── DocumentAnalysisTests.cs - Document analysis models
│ ├── EngineSettingsTests.cs - Engine configuration
│ └── ModelSettingsTests.cs - Model settings
├── Services/ (~41 tests)
│ ├── MockServiceTests.cs - Service mocks
│ └── ServiceInterfaceContractTests.cs - Interface contracts
└── Integration/ (~300 tests)
├── MemoryIntegrationTests.cs - Memory CRUD, search, context
├── EmbeddingIntegrationTests.cs - Embedding generation, similarity
├── RagIntegrationTests.cs - RAG document ingestion, retrieval
├── AdvancedRagTests.cs - Hybrid search, reranking
├── ConversationIntegrationTests.cs - Conversation and message handling
├── SearchIntegrationTests.cs - Message search, highlighting
├── LocalModelIntegrationTests.cs - Local model engine, streaming
├── RobustLocalModelTests.cs - Model state, download, memory
├── IngestionPipelineTests.cs - Document queue, processing
├── MemoryContextIntegrationTests.cs - Context generation, summarization
├── AiServiceIntegrationTests.cs - AI generation, summarization
├── EndToEndWorkflowTests.cs - Complete workflows
├── ErrorHandlingIntegrationTests.cs - Error scenarios, retry
├── CacheOptimizationTests.cs - Caching, batching, lazy loading
├── DocumentProcessingTests.cs - Chunking, analysis
├── LocalizationIntegrationTests.cs - Polish/English, formatting
├── SecurityIntegrationTests.cs - API keys, sanitization
├── ExportImportIntegrationTests.cs - Export, backup
├── OnboardingIntegrationTests.cs - First-run experience
├── ModelDownloadIntegrationTests.cs - Download progress, installation
├── ConcurrencyIntegrationTests.cs - Thread safety, rate limiting
├── PerformanceIntegrationTests.cs - Response time, throughput
└── OpenRouterIntegrationTests.cs - Real API integration tests
# Run all unit tests (excluding integration tests with real API)
dotnet test LLMClient.Tests --filter "Category!=Integration"
# Run all tests including integration tests
dotnet test LLMClient.Tests
# Run specific test category
dotnet test LLMClient.Tests --filter "FullyQualifiedName~Memory"
| Area | Tests | Description |
|---|---|---|
| Models | 163 | Data models, enums, properties |
| Memory | 23 | Memory CRUD, search, context building |
| Embeddings | 20 | Generation, similarity, model selection |
| RAG | 28 | Document ingestion, retrieval, hybrid search |
| Conversations | 15 | CRUD, messages, pagination |
| Search | 12 | Highlighting, Polish chars, async |
| Local Models | 28 | State management, download, inference |
| AI Service | 20 | Generation, streaming, summarization |
| Error Handling | 15 | Network, database, retry logic |
| Performance | 13 | Response time, memory, throughput |
| Security | 17 | API keys, sanitization, validation |
| Localization | 12 | Polish/English, date/number formats |
| Export | 9 | JSON, Markdown, backup/restore |
Some tests require an OpenRouter API key stored at ~/.llmclient/openrouter_api_key.txt:
# Create the file with your API key
echo "sk-or-your-api-key" > ~/.llmclient/openrouter_api_key.txt
These tests verify real API behavior including streaming, system prompts, and error handling.
This project is licensed under the MIT license.
LLMClient to zaawansowany, wieloplatformowy klient AI, stworzony w .NET MAUI, który umożliwia płynną interakcję z wiodącymi modelami językowymi (LLM), takimi jak GPT i Gemini. Aplikacja została zaprojektowana z myślą o bezpieczeństwie, wydajności i elastyczności, oferując lokalne przechowywanie danych z szyfrowaniem oraz natywną tokenizację tekstu.
tokenizersSklonuj repozytorium:
git clone https://github.com/DamianTarnowski/LLMClient.git
cd LLMClient
Zbuduj bibliotekę Rust:
cd TokenizerRust
cargo build --release
Spowoduje to umieszczenie pliku tokenizer_rust.dll (lub odpowiednika dla Twojego systemu) w folderze target/release. Projekt .NET MAUI jest skonfigurowany, aby automatycznie kopiować ten plik podczas budowania.
Otwórz i uruchom projekt .NET MAUI:
LLMClient.sln w programie Visual Studio 2022.Ustawienia (ikona koła zębatego).LLMClient posiada zaawansowany system pamięci, który automatycznie personalizuje interakcje z modelami AI:
Automatyczna Ekstraktacja: System automatycznie wydobywa informacje osobiste z konwersacji (imię, wiek, zawód, preferencje, itp.) przy użyciu:
Kontekst w Konwersacji: Każda konwersacja otrzymuje pełny kontekst pamięci (maksymalnie 30,000 znaków) w wiadomości systemowej, pozwalając modelom na spersonalizowane odpowiedzi.
Inteligentne Zarządzanie: Gdy pamięć przekroczy limit znaków:
Zarządzanie Pamięcią: Dostępny jest interfejs do przeglądania, edycji i kategoryzowania zapisanych wspomnień.
System automatycznie rozpoznaje i zapisuje:
Wszystkie dane pamięci są:
Projekt wykorzystuje wzorzec MVVM i jest podzielony na następujące warstwy:
LLMClient/: Główny projekt .NET MAUI.
Views/: Pliki XAML definiujące interfejs użytkownika.ViewModels/: Logika biznesowa i stan dla poszczególnych widoków.Models/: Modele danych (np. Conversation, Message).Services/: Serwisy odpowiedzialne za kluczowe funkcjonalności (np. AiService, DatabaseService, EmbeddingService, MemoryContextService, MemoryExtractionService).TokenizerRust/: Projekt Rust zawierający logikę tokenizacji, udostępnioną jako natywna biblioteka.SafeLocalModelWrapper opakowuje ILocalModelService, śledzi błędy i wymusza cooldownRobustLocalModelService, LlamaSharpLocalModelService wybierane przez EngineSettingsLocalModelDiagnosticService z okresowymi checkami i zapisem raportówNetworkAwareDownloadService z auto‑wznawianiem i świadomością łącznościMainPageViewModel subskrybuje zdarzenia MessagingCenter (LocalModelLoaded/LocalModelUnloaded, ModelsChanged), blokuje wybór chmury gdy lokalny jest aktywnyPlanowane ulepszenia znajdziesz tutaj: ROADMAP.md
LLMClient posiada kompleksowe pokrycie testami z ponad 500 automatycznymi testami obejmującymi wszystkie główne funkcjonalności.
LLMClient.Tests/
├── Models/ (~163 testy jednostkowe)
│ ├── ModelsTests.cs - Conversation, Message, Memory, AiModel
│ ├── EdgeCaseTests.cs - Przypadki brzegowe
│ ├── RagTraceTests.cs - Śledzenie RAG i timing
│ ├── ErrorReportTests.cs - Modele raportowania błędów
│ ├── EmbeddingModelTests.cs - Wybór modelu embeddingowego
│ ├── DocumentAnalysisTests.cs - Modele analizy dokumentów
│ ├── EngineSettingsTests.cs - Konfiguracja silnika
│ └── ModelSettingsTests.cs - Ustawienia modelu
├── Services/ (~41 testów)
│ ├── MockServiceTests.cs - Mocki serwisów
│ └── ServiceInterfaceContractTests.cs - Kontrakty interfejsów
└── Integration/ (~300 testów)
├── MemoryIntegrationTests.cs - Pamięć CRUD, wyszukiwanie
├── EmbeddingIntegrationTests.cs - Generowanie embeddingów
├── RagIntegrationTests.cs - Ingestion dokumentów RAG
├── AdvancedRagTests.cs - Hybrid search, reranking
├── ConversationIntegrationTests.cs - Obsługa konwersacji
├── SearchIntegrationTests.cs - Wyszukiwanie, highlighting
├── LocalModelIntegrationTests.cs - Silnik modeli lokalnych
├── RobustLocalModelTests.cs - Stan modelu, pobieranie
├── IngestionPipelineTests.cs - Kolejka dokumentów
├── MemoryContextIntegrationTests.cs - Generowanie kontekstu
├── AiServiceIntegrationTests.cs - Generowanie AI
├── EndToEndWorkflowTests.cs - Kompletne workflow
├── ErrorHandlingIntegrationTests.cs - Obsługa błędów
├── CacheOptimizationTests.cs - Cache, batching
├── DocumentProcessingTests.cs - Chunking, analiza
├── LocalizationIntegrationTests.cs - PL/EN, formatowanie
├── SecurityIntegrationTests.cs - API keys, sanityzacja
├── ExportImportIntegrationTests.cs - Eksport, backup
├── OnboardingIntegrationTests.cs - Pierwsze uruchomienie
├── ModelDownloadIntegrationTests.cs - Pobieranie modeli
├── ConcurrencyIntegrationTests.cs - Bezpieczeństwo wątków
├── PerformanceIntegrationTests.cs - Wydajność
└── OpenRouterIntegrationTests.cs - Testy z prawdziwym API
# Uruchom testy jednostkowe (bez testów API)
dotnet test LLMClient.Tests --filter "Category!=Integration"
# Uruchom wszystkie testy
dotnet test LLMClient.Tests
# Uruchom konkretną kategorię
dotnet test LLMClient.Tests --filter "FullyQualifiedName~Memory"
| Obszar | Testy | Opis |
|---|---|---|
| Modele | 163 | Modele danych, enumy, właściwości |
| Pamięć | 23 | CRUD, wyszukiwanie, budowanie kontekstu |
| Embeddingi | 20 | Generowanie, podobieństwo, wybór modelu |
| RAG | 28 | Ingestion, retrieval, hybrid search |
| Konwersacje | 15 | CRUD, wiadomości, paginacja |
| Wyszukiwanie | 12 | Highlighting, polskie znaki |
| Modele lokalne | 28 | Zarządzanie stanem, pobieranie |
| AI Service | 20 | Generowanie, streaming |
| Błędy | 15 | Sieć, baza danych, retry |
| Wydajność | 13 | Czas odpowiedzi, pamięć |
| Bezpieczeństwo | 17 | API keys, sanityzacja |
| Lokalizacja | 12 | PL/EN, formaty dat/liczb |
| Eksport | 9 | JSON, Markdown, backup |
Niektóre testy wymagają klucza API OpenRouter w ~/.llmclient/openrouter_api_key.txt:
# Utwórz plik z kluczem API
echo "sk-or-your-api-key" > ~/.llmclient/openrouter_api_key.txt
Ten projekt jest udostępniany na licencji MIT.
LLMClient offers two embedding models for semantic search and RAG (Retrieval-Augmented Generation):
| Model | Quality | Speed | Size | Min RAM | Best For |
|---|---|---|---|---|---|
| E5-Large Multilingual | 🎯 95% | 60% | ~2.2GB | 8GB | Desktop, high-quality search |
| Gemma 300M | 🚀 75% | 90% | ~1.2GB | 4GB | Mobile, resource-constrained devices |
The app automatically selects the best model based on device RAM:
Users can override this in Settings → Embedding Model.
Tests performed on Polish language pairs comparing semantic similarity scores:
| Pair | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Uwielbiam kodować" | 0.936 | 0.840 | +11% |
| "Kot śpi" ↔ "Kotek drzemie" | 0.944 | 0.614 | +54% |
| "Auto jedzie" ↔ "Samochód pędzi" | 0.952 | 0.682 | +40% |
| Average | 0.944 | 0.712 | +33% |
| Pair | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Pierogi są pyszne" | 0.829 | 0.441 | -47% |
| "Kot śpi" ↔ "Samolot leci" | 0.833 | 0.440 | -47% |
| "Auto jedzie" ↔ "Książka leży" | 0.788 | 0.393 | -50% |
| Average | 0.817 | 0.425 | -48% |
| Pair | E5-Large | Gemma 300M |
|---|---|---|
| PL: "Lubię programować" ↔ EN: "I like programming" | 0.966 | 0.701 |
| PL: "Kot jest zwierzęciem" ↔ EN: "A cat is an animal" | 0.920 | 0.794 |
| Average | 0.943 | 0.748 |
E5-Large excels at Polish semantic similarity - 33% higher scores for similar sentences, making it ideal for precise semantic search in Polish documents.
Gemma better discriminates unrelated content - 48% lower scores for different sentences means fewer false positives in search results.
E5-Large superior for cross-lingual - Critical for applications mixing Polish and English content.
Trade-off: E5-Large requires 2x more RAM but provides significantly better quality for Polish language semantic search. For mobile devices with <8GB RAM, Gemma offers acceptable quality with much better performance.
Both models use native Rust tokenizers for optimal performance:
How to use the local model
LLMClient oferuje dwa modele embeddingowe do wyszukiwania semantycznego i RAG:
| Model | Jakość | Szybkość | Rozmiar | Min RAM | Najlepszy dla |
|---|---|---|---|---|---|
| E5-Large Multilingual | 🎯 95% | 60% | ~2.2GB | 8GB | Desktop, wysokiej jakości wyszukiwanie |
| Gemma 300M | 🚀 75% | 90% | ~1.2GB | 4GB | Mobile, urządzenia z ograniczeniami |
Aplikacja automatycznie wybiera najlepszy model na podstawie RAM urządzenia:
Użytkownik może to zmienić w Ustawieniach → Model Embeddingowy.
Testy przeprowadzone na polskich parach zdań porównujące wyniki podobieństwa semantycznego:
| Para | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Uwielbiam kodować" | 0.936 | 0.840 | +11% |
| "Kot śpi" ↔ "Kotek drzemie" | 0.944 | 0.614 | +54% |
| "Auto jedzie" ↔ "Samochód pędzi" | 0.952 | 0.682 | +40% |
| Średnia | 0.944 | 0.712 | +33% |
| Para | E5-Large | Gemma 300M | Δ |
|---|---|---|---|
| "Lubię programować" ↔ "Pierogi są pyszne" | 0.829 | 0.441 | -47% |
| "Kot śpi" ↔ "Samolot leci" | 0.833 | 0.440 | -47% |
| "Auto jedzie" ↔ "Książka leży" | 0.788 | 0.393 | -50% |
| Średnia | 0.817 | 0.425 | -48% |
| Para | E5-Large | Gemma 300M |
|---|---|---|
| PL: "Lubię programować" ↔ EN: "I like programming" | 0.966 | 0.701 |
| PL: "Kot jest zwierzęciem" ↔ EN: "A cat is an animal" | 0.920 | 0.794 |
| Średnia | 0.943 | 0.748 |
E5-Large doskonały dla polskiego podobieństwa semantycznego - 33% wyższe wyniki dla podobnych zdań, idealny do precyzyjnego wyszukiwania semantycznego w polskich dokumentach.
Gemma lepiej rozróżnia niepowiązane treści - 48% niższe wyniki dla różnych zdań oznacza mniej fałszywych trafień w wynikach wyszukiwania.
E5-Large przewyższa w cross-lingual - Krytyczne dla aplikacji mieszających polskie i angielskie treści.
Kompromis: E5-Large wymaga 2x więcej RAM, ale zapewnia znacznie lepszą jakość dla polskiego wyszukiwania semantycznego. Dla urządzeń mobilnych z <8GB RAM, Gemma oferuje akceptowalną jakość przy znacznie lepszej wydajności.
Oba modele używają natywnych tokenizerów Rust dla optymalnej wydajności:
Jak korzystać z modelu lokalnego
C#
98.0%
Rust
1.4%