Here's the updated README.md incorporating the latest changes to the project, including GPU detection and the combined repository structure:
This project is a demonstration for Small Language Models (SLMs), presented as part of a lecture by Alon Fliess. It highlights how to load and interact with the Phi-3.5-mini-instruct-onnx model locally using C#. The demo dynamically determines whether to use GPU or CPU for model inference, ensuring optimal performance on the available hardware.
Model Directory Check:
Hardware Detection:
nvidia-smi command.Model and Tokenizer Initialization:
System Prompt Setup:
User Input:
CLH: Clears the conversation history and starts a new session.History Management:
Response Generation:
Output:
CloneGitRepository(string repoUrl, string destinationPath):
IsGpuAvailable():
nvidia-smi command.true if a GPU is detected.tokenizer.Encode(fullPrompt):
new GeneratorParams(model):
max_length).generator.ComputeLogits():
generator.GenerateNextToken():
tokenizer.Decode(newToken):
Clone this repository:
git clone <repository_url>
cd <repository_folder>
Restore dependencies:
dotnet restore
Build and run the project:
dotnet run
Interact with the AI assistant:
CLH command to clear history.Console Input:
Q: What is the capital of France?
Model Output:
Phi3.5: The capital of France is Paris.
The Phi-3.5-mini-instruct-onnx repository contains both GPU and CPU versions:
gpu/gpu-int4-awq-block-128.cpu_and_mobile/cpu-int4-awq-block-128-acc-level-4.gpu/gpu-int4-awq-block-128.cpu_and_mobile/cpu-int4-awq-block-128-acc-level-4.This project is designed to showcase the capabilities of Small Language Models and their efficiency in local AI inference tasks. It demonstrates how lightweight models can be effectively utilized in real-world scenarios. Feel free to explore and adapt the code for your own use cases.
C#
100.0%
Here's the updated README.md incorporating the latest changes to the project, including GPU detection and the combined repository structure:
This project is a demonstration for Small Language Models (SLMs), presented as part of a lecture by Alon Fliess. It highlights how to load and interact with the Phi-3.5-mini-instruct-onnx model locally using C#. The demo dynamically determines whether to use GPU or CPU for model inference, ensuring optimal performance on the available hardware.
Model Directory Check:
Hardware Detection:
nvidia-smi command.Model and Tokenizer Initialization:
System Prompt Setup:
User Input:
CLH: Clears the conversation history and starts a new session.History Management:
Response Generation:
Output:
CloneGitRepository(string repoUrl, string destinationPath):
IsGpuAvailable():
nvidia-smi command.true if a GPU is detected.tokenizer.Encode(fullPrompt):
new GeneratorParams(model):
max_length).generator.ComputeLogits():
generator.GenerateNextToken():
tokenizer.Decode(newToken):
Clone this repository:
git clone <repository_url>
cd <repository_folder>
Restore dependencies:
dotnet restore
Build and run the project:
dotnet run
Interact with the AI assistant:
CLH command to clear history.Console Input:
Q: What is the capital of France?
Model Output:
Phi3.5: The capital of France is Paris.
The Phi-3.5-mini-instruct-onnx repository contains both GPU and CPU versions:
gpu/gpu-int4-awq-block-128.cpu_and_mobile/cpu-int4-awq-block-128-acc-level-4.gpu/gpu-int4-awq-block-128.cpu_and_mobile/cpu-int4-awq-block-128-acc-level-4.This project is designed to showcase the capabilities of Small Language Models and their efficiency in local AI inference tasks. It demonstrates how lightweight models can be effectively utilized in real-world scenarios. Feel free to explore and adapt the code for your own use cases.
C#
100.0%