ashleyashok/fashion-knowledge-graph

5

stars

36

commits

Python

primary language

Aug 25, 2025

updated

README

Complete the Look: Fashion Recommendation System

Python License Code Style Architecture

A revolutionary fashion recommendation system that leverages knowledge graphs, computer vision, and social media intelligence to provide personalized "Complete the Look" suggestions. This system goes beyond traditional collaborative filtering by incorporating real-world fashion trends from social media, creating a dynamic knowledge graph that captures complex relationships between fashion items.

๐Ÿ“– Implementation Details & Technical Deep Dive

Want to understand how this system works under the hood?

๐Ÿ‘‰ Read our comprehensive technical blog post that explains:

  • The innovative social media-powered knowledge graph architecture
  • How we transform social media into a living style database
  • The dual-path search system using CLIP embeddings
  • Real-world performance metrics and business impact
  • The paradigm shift in fashion recommendation technology

๐Ÿ“‹ Quick Implementation Overview - High-level technical summary

The blog post provides detailed technical explanations, while this README focuses on practical usage and setup.

๐Ÿš€ Key Features

๐ŸŽจ Complete the Look Recommendations

  • Graph Traversal: Navigate the knowledge graph to find items frequently worn together
  • Trend-Aware: Prioritize recommendations based on current social media trends
  • Contextual Filtering: Consider season, occasion, and style preferences

๐Ÿ” Style Matching

  • Image Upload: Upload outfit images to find similar products in the catalog
  • Text Description: Describe your style preferences in natural language
  • Multi-Modal Search: Combine visual and textual similarity for accurate matching

๐Ÿ“Š Product Attribute Extraction

  • Unstructured Image Processing: Extract structured attributes from product spec sheets
  • LLM-Powered: Use GPT-4 for intelligent attribute extraction
  • Manual Override: Edit extracted attributes when needed

๐Ÿ”„ Social Media Integration

  • Real-time Trend Detection: Capture emerging fashion trends from social platforms
  • Co-occurrence Analysis: Identify items frequently worn together
  • Dynamic Updates: Continuously evolve the knowledge graph

๐Ÿ—๏ธ Architecture

The system consists of several interconnected components:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   Social Media  โ”‚    โ”‚   Catalog Data  โ”‚    โ”‚   User Input    โ”‚
โ”‚     Images      โ”‚    โ”‚                 โ”‚    โ”‚                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Image Processing Pipeline                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Segmentationโ”‚  โ”‚  Attribute  โ”‚  โ”‚  Embedding  โ”‚            โ”‚
โ”‚  โ”‚    Model    โ”‚  โ”‚ Extraction  โ”‚  โ”‚   Model     โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Knowledge Graph (Neo4j)                     โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚   Product   โ”‚  โ”‚  Relations  โ”‚  โ”‚  Attributes โ”‚            โ”‚
โ”‚  โ”‚    Nodes    โ”‚  โ”‚    Edges    โ”‚  โ”‚             โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                   Vector Database (Pinecone)                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚   Image     โ”‚  โ”‚   Style     โ”‚  โ”‚  Semantic   โ”‚            โ”‚
โ”‚  โ”‚ Embeddings  โ”‚  โ”‚ Embeddings  โ”‚  โ”‚   Search    โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Recommendation Engine                        โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Graph Query โ”‚  โ”‚ Vector      โ”‚  โ”‚ Hybrid      โ”‚            โ”‚
โ”‚  โ”‚ Traversal   โ”‚  โ”‚ Similarity  โ”‚  โ”‚ Ranking     โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Streamlit Interface                          โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Complete    โ”‚  โ”‚ Style       โ”‚  โ”‚ Attribute   โ”‚            โ”‚
โ”‚  โ”‚ the Look    โ”‚  โ”‚ Matching    โ”‚  โ”‚ Extraction  โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.11 or higher
  • Poetry for dependency management (recommended)
  • Neo4j database instance
  • Pinecone account for vector database
  • Azure OpenAI API key (for GPT-4)

Installation

  1. Clone the Repository

    git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
    cd fashion-knowledge-graph
    
  2. Install Dependencies

    poetry install
    
  3. Set Up Environment Variables

    cp .env.template .env
    # Edit .env with your API keys and database credentials
    
  4. Prepare Your Data

    • Place your catalog CSV file in output/data/catalog_combined.csv
    • Ensure it contains: product_id, image_path, category
  5. Initialize Databases

    # Create Pinecone indexes
    python scripts/setup_pinecone.py
    
    # Process catalog data
    python src/engine/process_catalog.py
    
  6. Launch the Application

    poetry run streamlit run app/main.py
    

Option 2: Using pip

  1. Clone the Repository

    git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
    cd fashion-knowledge-graph
    
  2. Install Dependencies

    pip install -r requirements.txt
    
  3. Set Up Environment Variables

    cp .env.template .env
    # Edit .env with your API keys and database credentials
    
  4. Launch the Application

    streamlit run app/main.py
    

โš™๏ธ Configuration

Environment Variables

Create a .env file with the following variables:

# Neo4j Database
NEO4J_URI=bolt://localhost:7687
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your_password

# Pinecone Vector Database
PINECONE_API_KEY=your_pinecone_api_key
PINECONE_HOST_IMAGE=your_image_index_host
PINECONE_HOST_STYLE=your_style_index_host

# Azure OpenAI
AZURE_OPENAI_API_KEY=your_azure_openai_key
AZURE_OPENAI_ENDPOINT=your_azure_openai_endpoint
AZURE_OPENAI_API_VERSION=2024-02-15-preview

# Model Configuration (optional - defaults provided)
SEGMENTATION_MODEL=sayeed99/segformer_b3_clothes
EMBEDDING_MODEL=Marqo/marqo-fashionCLIP
TEXT_EMBEDDING_MODEL=all-MiniLM-L6-v2

Model Configuration

The system uses several pre-trained models:

  • Segmentation: sayeed99/segformer_b3_clothes for clothing item detection
  • Image Embedding: Marqo/marqo-fashionCLIP for visual similarity
  • Text Embedding: all-MiniLM-L6-v2 for textual similarity
  • Attribute Extraction: Azure OpenAI GPT-4 for intelligent attribute extraction

๐Ÿ“– Usage Examples

Processing Catalog Data

Before using the system, you need to process your product catalog:

python src/engine/process_catalog.py

This will:

  • Extract attributes and embeddings from catalog images
  • Create product nodes in the Neo4j graph
  • Store embeddings in Pinecone vector database

Processing Social Media Images

To enrich the knowledge graph with real-world fashion combinations:

python src/engine/process_social_media_images.py

This will:

  • Analyze social media fashion images
  • Map detected items to catalog products
  • Create relationships based on co-occurrence patterns

Using the Streamlit Interface

The application provides three main features:

  1. Product Attribute Extraction

    • Upload product spec sheets or images
    • Extract structured attributes using AI
    • Edit attributes manually if needed
  2. Complete the Look

    • Select a product from the catalog
    • Get recommendations for complementary items
    • Filter by type, style, or occasion
  3. Style Matching

    • Upload outfit images or describe your style
    • Find similar products in the catalog
    • Get personalized recommendations

๐Ÿ—๏ธ Project Structure

complete-the-look/
โ”œโ”€โ”€ app/
โ”‚   โ””โ”€โ”€ main.py                   # Streamlit application
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ config/
โ”‚   โ”‚   โ””โ”€โ”€ settings.py           # Centralized configuration management
โ”‚   โ”œโ”€โ”€ models/
โ”‚   โ”‚   โ”œโ”€โ”€ base_model.py         # Abstract base classes for models
โ”‚   โ”‚   โ”œโ”€โ”€ segmentation_model.py # Image segmentation
โ”‚   โ”‚   โ”œโ”€โ”€ embedding_model.py    # Embedding generation
โ”‚   โ”‚   โ”œโ”€โ”€ model_manager.py      # Model lifecycle management
โ”‚   โ”‚   โ””โ”€โ”€ attribute_extraction_model.py
โ”‚   โ”œโ”€โ”€ database/
โ”‚   โ”‚   โ”œโ”€โ”€ graph_database.py     # Neo4j handler
โ”‚   โ”‚   โ””โ”€โ”€ vector_database.py    # Pinecone handler
โ”‚   โ”œโ”€โ”€ inference/
โ”‚   โ”‚   โ”œโ”€โ”€ recommender.py        # Recommendation engine
โ”‚   โ”‚   โ””โ”€โ”€ product_attributes.py
โ”‚   โ”œโ”€โ”€ engine/
โ”‚   โ”‚   โ”œโ”€โ”€ image_processor.py    # Core image processing orchestration
โ”‚   โ”‚   โ”œโ”€โ”€ process_catalog.py    # Catalog data processing
โ”‚   โ”‚   โ””โ”€โ”€ process_social_media_images.py  # Social media analysis
โ”‚   โ””โ”€โ”€ utils/
โ”‚       โ”œโ”€โ”€ models.py             # Data models and schemas
โ”‚       โ”œโ”€โ”€ prompts.py            # LLM prompts
โ”‚       โ””โ”€โ”€ tools.py              # Utility functions
โ”œโ”€โ”€ assets/
โ”‚   โ””โ”€โ”€ images/                   # Project images and diagrams
โ”œโ”€โ”€ output/
โ”‚   โ””โ”€โ”€ data/
โ”‚       โ””โ”€โ”€ catalog_combined.csv  # Catalog data file
โ”œโ”€โ”€ temp_images/                  # Temporary image storage
โ”œโ”€โ”€ scripts/                      # Setup and utility scripts
โ”œโ”€โ”€ docs/                         # Documentation
โ”œโ”€โ”€ tests/                        # Test files
โ”œโ”€โ”€ README.md                     # This file
โ”œโ”€โ”€ BLOG_POST.md                  # Technical blog post
โ”œโ”€โ”€ pyproject.toml               # Poetry configuration
โ””โ”€โ”€ .env                         # Environment variables

๐Ÿ“Š Performance Metrics

  • Recommendation Accuracy: 85% user satisfaction
  • Query Latency: <200ms for recommendation queries
  • Scalability: Handles 1M+ products and 10M+ relationships
  • Coverage: 90% of catalog items have meaningful relationships

๐Ÿงช Testing

Running Tests

# Install test dependencies
poetry install --with dev

# Run tests
poetry run pytest

# Run with coverage
poetry run pytest --cov=src

# Run specific test file
poetry run pytest tests/test_recommender.py

Code Quality Checks

# Format code
poetry run black src/ app/

# Lint code
poetry run flake8 src/ app/

# Type checking
poetry run mypy src/ app/

๐Ÿ”ง Development

Adding New Features

  1. Create a new module in the appropriate directory
  2. Follow the class-based pattern established in the codebase
  3. Add comprehensive docstrings and type hints
  4. Update the configuration if needed
  5. Add tests for new functionality

Code Style Guidelines

  • Follow PEP8 style guidelines
  • Use type hints for all function parameters and return values
  • Write comprehensive docstrings in Google style
  • Use meaningful variable and function names
  • Handle exceptions gracefully
  • Log important events and errors

๐Ÿ“Š Dataset

The fashion dataset contains 10,000+ images and is managed with Git LFS for efficient storage and retrieval.

Downloading the Dataset

# Clone with LFS files (recommended)
git lfs clone https://github.com/ashleyashok/fashion-knowledge-graph.git

# Or clone normally and pull LFS files
git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
cd fashion-knowledge-graph
git lfs pull

Dataset Structure

dataset/
โ”œโ”€โ”€ catalog_images/          # Product catalog images
โ”œโ”€โ”€ social_media_images/     # Social media fashion images
โ”œโ”€โ”€ test_images/            # Test and validation images
โ””โ”€โ”€ metadata/               # Image metadata and annotations

๐Ÿค Contributing

We welcome contributions! Please follow these guidelines:

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/amazing-feature
  3. Follow the code style: Use black, flake8, and mypy
  4. Add tests: Ensure new code is covered by tests
  5. Update documentation: Add docstrings and update README if needed
  6. Commit your changes: git commit -m 'Add amazing feature'
  7. Push to the branch: git push origin feature/amazing-feature
  8. Open a Pull Request

Development Setup

# Clone and setup
git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
cd fashion-knowledge-graph

# Install development dependencies
poetry install --with dev

# Setup pre-commit hooks
poetry run pre-commit install

# Run tests
poetry run pytest

๐Ÿ“ License

This project is licensed under the MIT License - see the LICENSE file for details.

๐Ÿ“š Additional Resources

๐Ÿง  Technical Deep Dive

๐Ÿ“Š Data & Management


โญ Star this repository if you find it useful!

Questions?

Contributors

ashleyashok

36 commits

ashleyashok/fashion-knowledge-graph

5

stars

36

commits

Python

primary language

Aug 25, 2025

updated

README

Complete the Look: Fashion Recommendation System

Python License Code Style Architecture

A revolutionary fashion recommendation system that leverages knowledge graphs, computer vision, and social media intelligence to provide personalized "Complete the Look" suggestions. This system goes beyond traditional collaborative filtering by incorporating real-world fashion trends from social media, creating a dynamic knowledge graph that captures complex relationships between fashion items.

๐Ÿ“– Implementation Details & Technical Deep Dive

Want to understand how this system works under the hood?

๐Ÿ‘‰ Read our comprehensive technical blog post that explains:

  • The innovative social media-powered knowledge graph architecture
  • How we transform social media into a living style database
  • The dual-path search system using CLIP embeddings
  • Real-world performance metrics and business impact
  • The paradigm shift in fashion recommendation technology

๐Ÿ“‹ Quick Implementation Overview - High-level technical summary

The blog post provides detailed technical explanations, while this README focuses on practical usage and setup.

๐Ÿš€ Key Features

๐ŸŽจ Complete the Look Recommendations

  • Graph Traversal: Navigate the knowledge graph to find items frequently worn together
  • Trend-Aware: Prioritize recommendations based on current social media trends
  • Contextual Filtering: Consider season, occasion, and style preferences

๐Ÿ” Style Matching

  • Image Upload: Upload outfit images to find similar products in the catalog
  • Text Description: Describe your style preferences in natural language
  • Multi-Modal Search: Combine visual and textual similarity for accurate matching

๐Ÿ“Š Product Attribute Extraction

  • Unstructured Image Processing: Extract structured attributes from product spec sheets
  • LLM-Powered: Use GPT-4 for intelligent attribute extraction
  • Manual Override: Edit extracted attributes when needed

๐Ÿ”„ Social Media Integration

  • Real-time Trend Detection: Capture emerging fashion trends from social platforms
  • Co-occurrence Analysis: Identify items frequently worn together
  • Dynamic Updates: Continuously evolve the knowledge graph

๐Ÿ—๏ธ Architecture

The system consists of several interconnected components:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   Social Media  โ”‚    โ”‚   Catalog Data  โ”‚    โ”‚   User Input    โ”‚
โ”‚     Images      โ”‚    โ”‚                 โ”‚    โ”‚                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Image Processing Pipeline                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Segmentationโ”‚  โ”‚  Attribute  โ”‚  โ”‚  Embedding  โ”‚            โ”‚
โ”‚  โ”‚    Model    โ”‚  โ”‚ Extraction  โ”‚  โ”‚   Model     โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Knowledge Graph (Neo4j)                     โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚   Product   โ”‚  โ”‚  Relations  โ”‚  โ”‚  Attributes โ”‚            โ”‚
โ”‚  โ”‚    Nodes    โ”‚  โ”‚    Edges    โ”‚  โ”‚             โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                   Vector Database (Pinecone)                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚   Image     โ”‚  โ”‚   Style     โ”‚  โ”‚  Semantic   โ”‚            โ”‚
โ”‚  โ”‚ Embeddings  โ”‚  โ”‚ Embeddings  โ”‚  โ”‚   Search    โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Recommendation Engine                        โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Graph Query โ”‚  โ”‚ Vector      โ”‚  โ”‚ Hybrid      โ”‚            โ”‚
โ”‚  โ”‚ Traversal   โ”‚  โ”‚ Similarity  โ”‚  โ”‚ Ranking     โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ–ผ                      โ–ผ                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Streamlit Interface                          โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚ Complete    โ”‚  โ”‚ Style       โ”‚  โ”‚ Attribute   โ”‚            โ”‚
โ”‚  โ”‚ the Look    โ”‚  โ”‚ Matching    โ”‚  โ”‚ Extraction  โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.11 or higher
  • Poetry for dependency management (recommended)
  • Neo4j database instance
  • Pinecone account for vector database
  • Azure OpenAI API key (for GPT-4)

Installation

  1. Clone the Repository

    git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
    cd fashion-knowledge-graph
    
  2. Install Dependencies

    poetry install
    
  3. Set Up Environment Variables

    cp .env.template .env
    # Edit .env with your API keys and database credentials
    
  4. Prepare Your Data

    • Place your catalog CSV file in output/data/catalog_combined.csv
    • Ensure it contains: product_id, image_path, category
  5. Initialize Databases

    # Create Pinecone indexes
    python scripts/setup_pinecone.py
    
    # Process catalog data
    python src/engine/process_catalog.py
    
  6. Launch the Application

    poetry run streamlit run app/main.py
    

Option 2: Using pip

  1. Clone the Repository

    git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
    cd fashion-knowledge-graph
    
  2. Install Dependencies

    pip install -r requirements.txt
    
  3. Set Up Environment Variables

    cp .env.template .env
    # Edit .env with your API keys and database credentials
    
  4. Launch the Application

    streamlit run app/main.py
    

โš™๏ธ Configuration

Environment Variables

Create a .env file with the following variables:

# Neo4j Database
NEO4J_URI=bolt://localhost:7687
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your_password

# Pinecone Vector Database
PINECONE_API_KEY=your_pinecone_api_key
PINECONE_HOST_IMAGE=your_image_index_host
PINECONE_HOST_STYLE=your_style_index_host

# Azure OpenAI
AZURE_OPENAI_API_KEY=your_azure_openai_key
AZURE_OPENAI_ENDPOINT=your_azure_openai_endpoint
AZURE_OPENAI_API_VERSION=2024-02-15-preview

# Model Configuration (optional - defaults provided)
SEGMENTATION_MODEL=sayeed99/segformer_b3_clothes
EMBEDDING_MODEL=Marqo/marqo-fashionCLIP
TEXT_EMBEDDING_MODEL=all-MiniLM-L6-v2

Model Configuration

The system uses several pre-trained models:

  • Segmentation: sayeed99/segformer_b3_clothes for clothing item detection
  • Image Embedding: Marqo/marqo-fashionCLIP for visual similarity
  • Text Embedding: all-MiniLM-L6-v2 for textual similarity
  • Attribute Extraction: Azure OpenAI GPT-4 for intelligent attribute extraction

๐Ÿ“– Usage Examples

Processing Catalog Data

Before using the system, you need to process your product catalog:

python src/engine/process_catalog.py

This will:

  • Extract attributes and embeddings from catalog images
  • Create product nodes in the Neo4j graph
  • Store embeddings in Pinecone vector database

Processing Social Media Images

To enrich the knowledge graph with real-world fashion combinations:

python src/engine/process_social_media_images.py

This will:

  • Analyze social media fashion images
  • Map detected items to catalog products
  • Create relationships based on co-occurrence patterns

Using the Streamlit Interface

The application provides three main features:

  1. Product Attribute Extraction

    • Upload product spec sheets or images
    • Extract structured attributes using AI
    • Edit attributes manually if needed
  2. Complete the Look

    • Select a product from the catalog
    • Get recommendations for complementary items
    • Filter by type, style, or occasion
  3. Style Matching

    • Upload outfit images or describe your style
    • Find similar products in the catalog
    • Get personalized recommendations

๐Ÿ—๏ธ Project Structure

complete-the-look/
โ”œโ”€โ”€ app/
โ”‚   โ””โ”€โ”€ main.py                   # Streamlit application
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ config/
โ”‚   โ”‚   โ””โ”€โ”€ settings.py           # Centralized configuration management
โ”‚   โ”œโ”€โ”€ models/
โ”‚   โ”‚   โ”œโ”€โ”€ base_model.py         # Abstract base classes for models
โ”‚   โ”‚   โ”œโ”€โ”€ segmentation_model.py # Image segmentation
โ”‚   โ”‚   โ”œโ”€โ”€ embedding_model.py    # Embedding generation
โ”‚   โ”‚   โ”œโ”€โ”€ model_manager.py      # Model lifecycle management
โ”‚   โ”‚   โ””โ”€โ”€ attribute_extraction_model.py
โ”‚   โ”œโ”€โ”€ database/
โ”‚   โ”‚   โ”œโ”€โ”€ graph_database.py     # Neo4j handler
โ”‚   โ”‚   โ””โ”€โ”€ vector_database.py    # Pinecone handler
โ”‚   โ”œโ”€โ”€ inference/
โ”‚   โ”‚   โ”œโ”€โ”€ recommender.py        # Recommendation engine
โ”‚   โ”‚   โ””โ”€โ”€ product_attributes.py
โ”‚   โ”œโ”€โ”€ engine/
โ”‚   โ”‚   โ”œโ”€โ”€ image_processor.py    # Core image processing orchestration
โ”‚   โ”‚   โ”œโ”€โ”€ process_catalog.py    # Catalog data processing
โ”‚   โ”‚   โ””โ”€โ”€ process_social_media_images.py  # Social media analysis
โ”‚   โ””โ”€โ”€ utils/
โ”‚       โ”œโ”€โ”€ models.py             # Data models and schemas
โ”‚       โ”œโ”€โ”€ prompts.py            # LLM prompts
โ”‚       โ””โ”€โ”€ tools.py              # Utility functions
โ”œโ”€โ”€ assets/
โ”‚   โ””โ”€โ”€ images/                   # Project images and diagrams
โ”œโ”€โ”€ output/
โ”‚   โ””โ”€โ”€ data/
โ”‚       โ””โ”€โ”€ catalog_combined.csv  # Catalog data file
โ”œโ”€โ”€ temp_images/                  # Temporary image storage
โ”œโ”€โ”€ scripts/                      # Setup and utility scripts
โ”œโ”€โ”€ docs/                         # Documentation
โ”œโ”€โ”€ tests/                        # Test files
โ”œโ”€โ”€ README.md                     # This file
โ”œโ”€โ”€ BLOG_POST.md                  # Technical blog post
โ”œโ”€โ”€ pyproject.toml               # Poetry configuration
โ””โ”€โ”€ .env                         # Environment variables

๐Ÿ“Š Performance Metrics

  • Recommendation Accuracy: 85% user satisfaction
  • Query Latency: <200ms for recommendation queries
  • Scalability: Handles 1M+ products and 10M+ relationships
  • Coverage: 90% of catalog items have meaningful relationships

๐Ÿงช Testing

Running Tests

# Install test dependencies
poetry install --with dev

# Run tests
poetry run pytest

# Run with coverage
poetry run pytest --cov=src

# Run specific test file
poetry run pytest tests/test_recommender.py

Code Quality Checks

# Format code
poetry run black src/ app/

# Lint code
poetry run flake8 src/ app/

# Type checking
poetry run mypy src/ app/

๐Ÿ”ง Development

Adding New Features

  1. Create a new module in the appropriate directory
  2. Follow the class-based pattern established in the codebase
  3. Add comprehensive docstrings and type hints
  4. Update the configuration if needed
  5. Add tests for new functionality

Code Style Guidelines

  • Follow PEP8 style guidelines
  • Use type hints for all function parameters and return values
  • Write comprehensive docstrings in Google style
  • Use meaningful variable and function names
  • Handle exceptions gracefully
  • Log important events and errors

๐Ÿ“Š Dataset

The fashion dataset contains 10,000+ images and is managed with Git LFS for efficient storage and retrieval.

Downloading the Dataset

# Clone with LFS files (recommended)
git lfs clone https://github.com/ashleyashok/fashion-knowledge-graph.git

# Or clone normally and pull LFS files
git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
cd fashion-knowledge-graph
git lfs pull

Dataset Structure

dataset/
โ”œโ”€โ”€ catalog_images/          # Product catalog images
โ”œโ”€โ”€ social_media_images/     # Social media fashion images
โ”œโ”€โ”€ test_images/            # Test and validation images
โ””โ”€โ”€ metadata/               # Image metadata and annotations

๐Ÿค Contributing

We welcome contributions! Please follow these guidelines:

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/amazing-feature
  3. Follow the code style: Use black, flake8, and mypy
  4. Add tests: Ensure new code is covered by tests
  5. Update documentation: Add docstrings and update README if needed
  6. Commit your changes: git commit -m 'Add amazing feature'
  7. Push to the branch: git push origin feature/amazing-feature
  8. Open a Pull Request

Development Setup

# Clone and setup
git clone https://github.com/ashleyashok/fashion-knowledge-graph.git
cd fashion-knowledge-graph

# Install development dependencies
poetry install --with dev

# Setup pre-commit hooks
poetry run pre-commit install

# Run tests
poetry run pytest

๐Ÿ“ License

This project is licensed under the MIT License - see the LICENSE file for details.

๐Ÿ“š Additional Resources

๐Ÿง  Technical Deep Dive

๐Ÿ“Š Data & Management


โญ Star this repository if you find it useful!

Questions?

Contributors

ashleyashok

36 commits

Languages

Python

100.0%