BLIP Image Captioning - Arabic (Flickr8k Arabic)
1
12 commits
1 linked in READMEs
updated May 10, 2025
This model is a fine-tuned version of Salesforce/blip-image-captioning-large, adapted for image captioning in Arabic using the Flickr8K Arabic dataset. It takes an input image and generates a relevant caption in Arabic, describing the image content.
from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image
import torch
import matplotlib.pyplot as plt
# Load model and processor
processor = BlipProcessor.from_pretrained("omarsabri8756/blip-Arabic-flickr-8k")
model = BlipForConditionalGeneration.from_pretrained("omarsabri8756/blip-Arabic-flickr-8k")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device)
# Load an image from local path
image_path = "path/to/your/image.jpg"
image = Image.open(image_path).convert("RGB")
# Show image
plt.imshow(image)
plt.axis('off')
plt.title("Input Image")
plt.show()
# Generate enhanced Arabic caption with better parameters
model.eval()
with torch.no_grad():
pixel_values = processor(images=image, return_tensors="pt").pixel_values.to(device)
generated_output = model.generate(
pixel_values=pixel_values,
max_length=75,
min_length=20,
num_beams=5,
repetition_penalty=1.5,
length_penalty=1.0,
no_repeat_ngram_size=3,
early_stopping=True
)
caption = processor.batch_decode(generated_output, skip_special_tokens=True)[0]
print(caption) # Prints Arabic caption
This model was fine-tuned on the Flickr8k Arabic dataset, which consists of 8,000 images, each with 4 reference Arabic captions. The dataset provides a diverse collection of everyday scenes and activities described in Modern Standard Arabic.
The model was fine-tuned from the original BLIP model by adapting its language generation capabilities to Arabic text.
The model was evaluated on the Flickr8k Arabic test split, which contains 1,000 images with 4 reference captions each.
The model performs well on common scenes and activities, generating grammatically correct and contextually appropriate Arabic captions. Performance decreases slightly for unusual scenes or culturally specific contexts not well-represented in the training data.
BLIP Image Captioning - Arabic (Flickr8k Arabic)
1
12 commits
1 linked in READMEs
updated May 10, 2025
This model is a fine-tuned version of Salesforce/blip-image-captioning-large, adapted for image captioning in Arabic using the Flickr8K Arabic dataset. It takes an input image and generates a relevant caption in Arabic, describing the image content.
from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image
import torch
import matplotlib.pyplot as plt
# Load model and processor
processor = BlipProcessor.from_pretrained("omarsabri8756/blip-Arabic-flickr-8k")
model = BlipForConditionalGeneration.from_pretrained("omarsabri8756/blip-Arabic-flickr-8k")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device)
# Load an image from local path
image_path = "path/to/your/image.jpg"
image = Image.open(image_path).convert("RGB")
# Show image
plt.imshow(image)
plt.axis('off')
plt.title("Input Image")
plt.show()
# Generate enhanced Arabic caption with better parameters
model.eval()
with torch.no_grad():
pixel_values = processor(images=image, return_tensors="pt").pixel_values.to(device)
generated_output = model.generate(
pixel_values=pixel_values,
max_length=75,
min_length=20,
num_beams=5,
repetition_penalty=1.5,
length_penalty=1.0,
no_repeat_ngram_size=3,
early_stopping=True
)
caption = processor.batch_decode(generated_output, skip_special_tokens=True)[0]
print(caption) # Prints Arabic caption
This model was fine-tuned on the Flickr8k Arabic dataset, which consists of 8,000 images, each with 4 reference Arabic captions. The dataset provides a diverse collection of everyday scenes and activities described in Modern Standard Arabic.
The model was fine-tuned from the original BLIP model by adapting its language generation capabilities to Arabic text.
The model was evaluated on the Flickr8k Arabic test split, which contains 1,000 images with 4 reference captions each.
The model performs well on common scenes and activities, generating grammatically correct and contextually appropriate Arabic captions. Performance decreases slightly for unusual scenes or culturally specific contexts not well-represented in the training data.