cartesia-ai/Llamba-3B

Model

Llamba Models

3

5 commits

1 linked in READMEs

updated Mar 15, 2025

See the code

README

Llamba Models

The Llamba models are part of Cartesia's Edge library, designed for efficient, high-performance machine learning applications.

For more details, refer to the paper.


Usage

Llamba on PyTorch

To use Llamba with PyTorch:

  1. Install the required package:
pip install --no-binary :all: cartesia-pytorch
  1. Load and run the model
from transformers import AutoTokenizer
from cartesia_pytorch.Llamba.llamba import LlambaLMHeadModel

model = LlambaLMHeadModel.from_pretrained("cartesia-ai/Llamba-3B", strict=True).to('cuda')
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-3B")
input_ids = tokenizer("Hello, my name is", return_tensors="pt").input_ids
input_ids = input_ids.to('cuda')
output = model.generate(input_ids, max_length=100)[0]
print(tokenizer.decode(output, skip_special_tokens=True))

Llamba on MLX

To run Llamba with the Metal framework see cartesia-metal


Evaluations

The Llamba models have been evaluated on multiple standard benchmarks, demonstrating efficiency gains while maintaining strong performance. Below are the results:

ModelARC-C (0-shot)ARC-C (25-shot)ARC-E (0-shot)ARC-E (25-shot)PIQA (0-shot)PIQA (10-shot)WG (0-shot)WG (5-shot)
Llamba-1B37.241.869.571.274.074.360.658.1
Llamba-3B48.553.079.081.178.679.570.472.4
Llamba-8B54.660.082.585.880.981.573.376.9
ModelHS (0-shot)HS (10-shot)LMB (0-shot)LMB (10-shot)MMLU (0-shot)MMLU (5-shot)OBQA (0-shot)OBQA (10-shot)
Llamba-1B61.260.248.439.038.031.337.038.0
Llamba-3B73.874.365.860.052.750.342.842.8
Llamba-8B77.678.769.465.061.060.043.445.8

More details on model performance, benchmarks, and evaluation metrics can be found in the paper.

cartesia
cartesia-pytorch
distillation
edge
llamba
Llamba
recurrent-models
safetensors

cartesia-ai/Llamba-3B

Model

Llamba Models

3

5 commits

1 linked in READMEs

updated Mar 15, 2025

See the code

README

Llamba Models

The Llamba models are part of Cartesia's Edge library, designed for efficient, high-performance machine learning applications.

For more details, refer to the paper.


Usage

Llamba on PyTorch

To use Llamba with PyTorch:

  1. Install the required package:
pip install --no-binary :all: cartesia-pytorch
  1. Load and run the model
from transformers import AutoTokenizer
from cartesia_pytorch.Llamba.llamba import LlambaLMHeadModel

model = LlambaLMHeadModel.from_pretrained("cartesia-ai/Llamba-3B", strict=True).to('cuda')
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-3B")
input_ids = tokenizer("Hello, my name is", return_tensors="pt").input_ids
input_ids = input_ids.to('cuda')
output = model.generate(input_ids, max_length=100)[0]
print(tokenizer.decode(output, skip_special_tokens=True))

Llamba on MLX

To run Llamba with the Metal framework see cartesia-metal


Evaluations

The Llamba models have been evaluated on multiple standard benchmarks, demonstrating efficiency gains while maintaining strong performance. Below are the results:

ModelARC-C (0-shot)ARC-C (25-shot)ARC-E (0-shot)ARC-E (25-shot)PIQA (0-shot)PIQA (10-shot)WG (0-shot)WG (5-shot)
Llamba-1B37.241.869.571.274.074.360.658.1
Llamba-3B48.553.079.081.178.679.570.472.4
Llamba-8B54.660.082.585.880.981.573.376.9
ModelHS (0-shot)HS (10-shot)LMB (0-shot)LMB (10-shot)MMLU (0-shot)MMLU (5-shot)OBQA (0-shot)OBQA (10-shot)
Llamba-1B61.260.248.439.038.031.337.038.0
Llamba-3B73.874.365.860.052.750.342.842.8
Llamba-8B77.678.769.465.061.060.043.445.8

More details on model performance, benchmarks, and evaluation metrics can be found in the paper.

cartesia
cartesia-pytorch
distillation
edge
llamba
Llamba
recurrent-models
safetensors