Each experiment (e.g., lexical content evaluation on chatbot arena evaluation) is organized into a separate folder. Each folder (roughly) contains the following subfolders:
data: Data used in the experimentnotebooks: Jupyter notebooks for data processing and analysisexperiments: Output cache files from running scriptsscripts: Python scripts for running the experimentslexical-content-chatbot-arena: Lexical content evaluation on chatbot arena evaluationspeech-quality: Speech quality assessment, Mean Opinion score (MOS) score prediction (naturalness, intelligibility, etc)paralinguistic: Paralinguistic evaluation, including data synthesized from ElevenLabs and GPT-4o-Audio (based on content from ChatbotArena)advanced-voice-gen-task-v1: Data preparation for advanced voice generation task v1 (the first version of SpeakBench)eval-leaderboard: Run inference for existing speech-in speech-out systems on the data from advanced-voice-gen-task-v1, and evaluate the outputs in a pairwise setup (like AlpacaEval) using AudioLLM-as-a-Judge.kokoroTTS: Scripts to run kokoroTTS for various experimentslegacy-exp: Legacy code for previous experiments70 commits
Jupyter Notebook
97.1%
Python
2.9%
Each experiment (e.g., lexical content evaluation on chatbot arena evaluation) is organized into a separate folder. Each folder (roughly) contains the following subfolders:
data: Data used in the experimentnotebooks: Jupyter notebooks for data processing and analysisexperiments: Output cache files from running scriptsscripts: Python scripts for running the experimentslexical-content-chatbot-arena: Lexical content evaluation on chatbot arena evaluationspeech-quality: Speech quality assessment, Mean Opinion score (MOS) score prediction (naturalness, intelligibility, etc)paralinguistic: Paralinguistic evaluation, including data synthesized from ElevenLabs and GPT-4o-Audio (based on content from ChatbotArena)advanced-voice-gen-task-v1: Data preparation for advanced voice generation task v1 (the first version of SpeakBench)eval-leaderboard: Run inference for existing speech-in speech-out systems on the data from advanced-voice-gen-task-v1, and evaluate the outputs in a pairwise setup (like AlpacaEval) using AudioLLM-as-a-Judge.kokoroTTS: Scripts to run kokoroTTS for various experimentslegacy-exp: Legacy code for previous experiments70 commits
Jupyter Notebook
97.1%
Python
2.9%