This package is a collection of datasets for evaluating AI Models in the context of Home Assistant.
See the codeThis package is a collection of datasets for evaluating AI Models in the context of Home Assistant. The overall approach is:
graph LR;
A[Synthetic Data Generation]
B[Dataset]
C[Collect Model Outputs]
D[Synthetic Home]
F[Evaluate Results]
G[Visualize Results]
H[OpenAI]
I[Conversation Agent]
J[Local LM]
K[Conversation Agent]
L[Google]
M[Conversation Agent]
A --> B
B --> D
D --> C
C --> F
F --> G
H --> I
J --> K
L --> M
I --> C
K --> C
M --> C
I --> D
K --> D
M --> D
See the datasets README for details on the available datasets including Home descriptions, Area descriptions, Device descriptions and summaries that can be performed on a home.
The device level datasets are defined using the Synthetic Home format including its device registry of synthetic devices.
See the generation README for more details on how synthetic data generation using LLMs works. The data is generated from a small amount of seed example data and a prompt, then is persisted.
The synthetic data generation is run with Jupyter notebooks.
classDiagram
direction LR
Home <|-- Area
Area <|-- Device
Device <|-- EntityStates
class Home{
+String name
+String country_code
+String location
+String type
}
class Area {
+String name
}
class Device {
+String name
+String device_type
+String model
+String mfg
+String sw_version
}
class EntityState {
+String state
}
You can use the generated synthetic data in Home Assistat and with integrated conversation agents to produce outputs for evaluation.
Model evaluation is currently performed with pytest, Synthetic Home, and any conversation agent (Open AI, Google, custom components, etc)
See [docs/eval.md] for instructions on how run an evaluation and update the leaderboard.
The most commonly used evaluation is for the Home Assistant conversation agent actions for integrating with the assist pipeline. See the following dataset directories for more information on running an evaluation:
Models are configured in models/.
We have an experimental dataset for zero show blueprint and automation creation, set up in a similar in style to a Software Engineer Benchmark. Each record contains a README with a description of the problem and expected results and a test eval that loads the blueprints generated by a model and exercises to verify if the solution is correct. See the dataset directory for more information:
There are additional datasets for human evaluation of summarization tasks. These were the initial use case for this repo. It works something like this:
These can be used for human evaluation to determine the model quality. In this phase, we take the model outputs from a human rater and use them for evaluation.
Human rater (me) scores the result quality:
See the script/ directory for more details on preparing the data for human eval procedure using Doccano.
398 followers · starred Aug 2024
104 followers · starred Oct 2024
48 followers · starred Feb 2025
27 followers · starred Sep 2025
Jupyter Notebook
67.1%
Python
31.6%
Shell
1.2%
This package is a collection of datasets for evaluating AI Models in the context of Home Assistant.
See the codeThis package is a collection of datasets for evaluating AI Models in the context of Home Assistant. The overall approach is:
graph LR;
A[Synthetic Data Generation]
B[Dataset]
C[Collect Model Outputs]
D[Synthetic Home]
F[Evaluate Results]
G[Visualize Results]
H[OpenAI]
I[Conversation Agent]
J[Local LM]
K[Conversation Agent]
L[Google]
M[Conversation Agent]
A --> B
B --> D
D --> C
C --> F
F --> G
H --> I
J --> K
L --> M
I --> C
K --> C
M --> C
I --> D
K --> D
M --> D
See the datasets README for details on the available datasets including Home descriptions, Area descriptions, Device descriptions and summaries that can be performed on a home.
The device level datasets are defined using the Synthetic Home format including its device registry of synthetic devices.
See the generation README for more details on how synthetic data generation using LLMs works. The data is generated from a small amount of seed example data and a prompt, then is persisted.
The synthetic data generation is run with Jupyter notebooks.
classDiagram
direction LR
Home <|-- Area
Area <|-- Device
Device <|-- EntityStates
class Home{
+String name
+String country_code
+String location
+String type
}
class Area {
+String name
}
class Device {
+String name
+String device_type
+String model
+String mfg
+String sw_version
}
class EntityState {
+String state
}
You can use the generated synthetic data in Home Assistat and with integrated conversation agents to produce outputs for evaluation.
Model evaluation is currently performed with pytest, Synthetic Home, and any conversation agent (Open AI, Google, custom components, etc)
See [docs/eval.md] for instructions on how run an evaluation and update the leaderboard.
The most commonly used evaluation is for the Home Assistant conversation agent actions for integrating with the assist pipeline. See the following dataset directories for more information on running an evaluation:
Models are configured in models/.
We have an experimental dataset for zero show blueprint and automation creation, set up in a similar in style to a Software Engineer Benchmark. Each record contains a README with a description of the problem and expected results and a test eval that loads the blueprints generated by a model and exercises to verify if the solution is correct. See the dataset directory for more information:
There are additional datasets for human evaluation of summarization tasks. These were the initial use case for this repo. It works something like this:
These can be used for human evaluation to determine the model quality. In this phase, we take the model outputs from a human rater and use them for evaluation.
Human rater (me) scores the result quality:
See the script/ directory for more details on preparing the data for human eval procedure using Doccano.
398 followers · starred Aug 2024
104 followers · starred Oct 2024
48 followers · starred Feb 2025
27 followers · starred Sep 2025
Jupyter Notebook
67.1%
Python
31.6%
Shell
1.2%