Repo for my Tensorflow/Keras CV experiments. Mostly revolving around the Danbooru20xx dataset
Python
128
92 commits
updated Nov 6, 2023
Repo for my Tensorflow/Keras CV experiments. Mostly revolving around the Danbooru20xx dataset
Framework: TF/Keras 2.10
Training SQLite DB built using fire-egg's tools: https://github.com/fire-eggs/Danbooru2019
Currently training on Danbooru2021, 512px SFW subset (sans the rating:q images that had been included in the 2022-01-21 release of the dataset)
Anonymous, The Danbooru Community, & Gwern Branwen; “Danbooru2021: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset”, 2022-01-21. Web. Accessed 2022-01-28 https://www.gwern.net/Danbooru2021
21/10/2022:
New release today, after a lot of time, many failed experiments, some successes and a whopping final 200 epochs and 24 days of training.
TPU time courtesy of TRC. Thanks!
ViT in particular was a pain, turns out it is quite sensitive to learning rate. I'm still not 100% sure I nailed it, but right now it works(TM).
While it might look like ConvNext does better, the whole story is a bit more interesting.
Full results below:
All 5500 tags:
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3611 | 0.7044 | 0.7044 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3673 | 0.6941 | 0.6941 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3663 | 0.6918 | 0.6918 |
All tags, starting from #2380 and below (sorted by most to least popular):
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3611 | 0.6107 | 0.5588 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3673 | 0.5932 | 0.5425 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3663 | 0.5993 | 0.5529 |
General tags (category 0), no character or series tags:
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3618 | 0.6878 | 0.6878 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3682 | 0.6774 | 0.6774 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3672 | 0.6748 | 0.6748 |
General tags (category 0), no character or series tags, starting from #2000 and below (sorted by most to least popular):
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3618 | 0.4515 | 0.3976 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3682 | 0.4320 | 0.3804 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3672 | 0.4416 | 0.3936 |
The numbers are obtained using tools/analyze_metrics.py to first find the point where P ≈ R, then using that threshold to check what scores I get on the less popular tags.
ViT blazes past ConvNext when it comes to rarer tags, so you might want to consider that when choosing what model to use.
Personally, I ensemble them if I don't have time constraints. That would be ensenble_october in the tables above. Quite some gains.
Next, I'll be finetuning at least one of these models on the latest tags, and adding NSFW images and tags to the training set, so that it can be used in tandem with Waifu Diffusion.
21/04/2022:
Checkpointing sweeps and conclusions so far:
05/04/2022:
So, I trained a bunch of ConvNexts in the past month.
Learned a few things too:
04/03/2022:
It is done! As the commit says, the first round of experiments on NFNets is finally over.
W&B reports this whole thing took about 112 days of compute time for the main runs, plus 11 days for experiments on higher image resolutions, and likely some more days on top accounting for failed experiments.
All of the experiments were run on a fleet of TPUv2 and TPUv3 VMs offered for free by the TRC program. Thanks!
Few takeaways, going over the data:
Now a word about the validation dataset:
This means the validation metrics are likely overoptimistic.
The first, obvious step would be to take the original images modulo 0900-0949, which I have reserved for this exact scenario, remove duplicates and "plausibly similar" images (say, maybe using embeddings or a very low threshold with a perceptual hash), preprocess them in a more uniform way, and use the output as a final test dataset.
18/02/2022:
So far I'm incredibly pleased with the results of adding ECA to Lx+SiLU networks.
At the meager cost of ~120 more parameters (give or take, depending on network depth) it is extremely effective at increasing network capacity.
On the other hand, I seem to have hit a wall with {L2,L1}+SiLU+ECA.
They overfit by epoch ~70 and ~85 out of 100, respectively.
I tried increasing MixUp to 0.3. This slowed down overfitting, but even then, by the end of training, the checkpoints with the best validation loss didn't display improved metrics over their MixUp 0.2 counterparts.
Something that DID work was finetuning while increasing the image size from 320 to 384.
I also tried a handful of learning rate schedules. In the end the one that worked best was starting off with max_learning_rate (0.1), no warmup, and letting cosine annealing do its thing over the course of 10 epochs.
06/02/2022:
Great news crew! TRC allowed me to use a bunch of TPUs!
To make better use of this amount of compute I had to overhaul a number of components, so a bunch of things are likely to have fallen to bitrot in the process.
I can only guarantee NFNet can work pretty much as before with the right arguments.
NFResNet changes should have left it retrocompatible with the previous version.
ResNet has been streamlined to be mostly in line with the Bag-of-Tricks paper (arXiv:1812.01187) with the exception of the stem. It is not compatible with the previous version of the code.
The training labels have been included in the 2021_0000_0899 folder for convenience.
The list of files used for training is going to be uploaded as a GitHub Release.
Now for some numbers:
compared to my previous best run, the one that resulted in NFNetL1V1-100-0.57141:
And it's all thanks to the folks at TRC, so shout out to them!
I currently have a few runs in progress across a couple of dimensions:
Once the experiments are over, the plan is to select the network definitions that lay on the Pareto curve between throughput and F1 score and release the trained weights.
One last thing.
I'd like to call your attention to the tools/cleanlab_stuff.py script.
It reads two files: one with the binarized labels from the database, the other with the predicted probabilities.
It then uses the cleanlab package to estimate whether if an image in a set could be missing a given label. At the end it stores its conclusions in a json file.
This file could, potentially, be used in some tool to assist human intervention to add the missing tags.
Python
100.0%
Repo for my Tensorflow/Keras CV experiments. Mostly revolving around the Danbooru20xx dataset
Python
128
92 commits
updated Nov 6, 2023
Repo for my Tensorflow/Keras CV experiments. Mostly revolving around the Danbooru20xx dataset
Framework: TF/Keras 2.10
Training SQLite DB built using fire-egg's tools: https://github.com/fire-eggs/Danbooru2019
Currently training on Danbooru2021, 512px SFW subset (sans the rating:q images that had been included in the 2022-01-21 release of the dataset)
Anonymous, The Danbooru Community, & Gwern Branwen; “Danbooru2021: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset”, 2022-01-21. Web. Accessed 2022-01-28 https://www.gwern.net/Danbooru2021
21/10/2022:
New release today, after a lot of time, many failed experiments, some successes and a whopping final 200 epochs and 24 days of training.
TPU time courtesy of TRC. Thanks!
ViT in particular was a pain, turns out it is quite sensitive to learning rate. I'm still not 100% sure I nailed it, but right now it works(TM).
While it might look like ConvNext does better, the whole story is a bit more interesting.
Full results below:
All 5500 tags:
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3611 | 0.7044 | 0.7044 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3673 | 0.6941 | 0.6941 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3663 | 0.6918 | 0.6918 |
All tags, starting from #2380 and below (sorted by most to least popular):
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3611 | 0.6107 | 0.5588 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3673 | 0.5932 | 0.5425 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3663 | 0.5993 | 0.5529 |
General tags (category 0), no character or series tags:
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3618 | 0.6878 | 0.6878 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3682 | 0.6774 | 0.6774 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3672 | 0.6748 | 0.6748 |
General tags (category 0), no character or series tags, starting from #2000 and below (sorted by most to least popular):
| run_name | definition_name | params_human | image_size | thres | F1 | F2 |
|---|---|---|---|---|---|---|
| ensenble_october | - | - | 448 | 0.3618 | 0.4515 | 0.3976 |
| ConvNextBV1_09_25_2022_05h13m55s | B | 93.2M | 448 | 0.3682 | 0.4320 | 0.3804 |
| ViTB16_09_25_2022_04h53m38s | B16 | 90.5M | 448 | 0.3672 | 0.4416 | 0.3936 |
The numbers are obtained using tools/analyze_metrics.py to first find the point where P ≈ R, then using that threshold to check what scores I get on the less popular tags.
ViT blazes past ConvNext when it comes to rarer tags, so you might want to consider that when choosing what model to use.
Personally, I ensemble them if I don't have time constraints. That would be ensenble_october in the tables above. Quite some gains.
Next, I'll be finetuning at least one of these models on the latest tags, and adding NSFW images and tags to the training set, so that it can be used in tandem with Waifu Diffusion.
21/04/2022:
Checkpointing sweeps and conclusions so far:
05/04/2022:
So, I trained a bunch of ConvNexts in the past month.
Learned a few things too:
04/03/2022:
It is done! As the commit says, the first round of experiments on NFNets is finally over.
W&B reports this whole thing took about 112 days of compute time for the main runs, plus 11 days for experiments on higher image resolutions, and likely some more days on top accounting for failed experiments.
All of the experiments were run on a fleet of TPUv2 and TPUv3 VMs offered for free by the TRC program. Thanks!
Few takeaways, going over the data:
Now a word about the validation dataset:
This means the validation metrics are likely overoptimistic.
The first, obvious step would be to take the original images modulo 0900-0949, which I have reserved for this exact scenario, remove duplicates and "plausibly similar" images (say, maybe using embeddings or a very low threshold with a perceptual hash), preprocess them in a more uniform way, and use the output as a final test dataset.
18/02/2022:
So far I'm incredibly pleased with the results of adding ECA to Lx+SiLU networks.
At the meager cost of ~120 more parameters (give or take, depending on network depth) it is extremely effective at increasing network capacity.
On the other hand, I seem to have hit a wall with {L2,L1}+SiLU+ECA.
They overfit by epoch ~70 and ~85 out of 100, respectively.
I tried increasing MixUp to 0.3. This slowed down overfitting, but even then, by the end of training, the checkpoints with the best validation loss didn't display improved metrics over their MixUp 0.2 counterparts.
Something that DID work was finetuning while increasing the image size from 320 to 384.
I also tried a handful of learning rate schedules. In the end the one that worked best was starting off with max_learning_rate (0.1), no warmup, and letting cosine annealing do its thing over the course of 10 epochs.
06/02/2022:
Great news crew! TRC allowed me to use a bunch of TPUs!
To make better use of this amount of compute I had to overhaul a number of components, so a bunch of things are likely to have fallen to bitrot in the process.
I can only guarantee NFNet can work pretty much as before with the right arguments.
NFResNet changes should have left it retrocompatible with the previous version.
ResNet has been streamlined to be mostly in line with the Bag-of-Tricks paper (arXiv:1812.01187) with the exception of the stem. It is not compatible with the previous version of the code.
The training labels have been included in the 2021_0000_0899 folder for convenience.
The list of files used for training is going to be uploaded as a GitHub Release.
Now for some numbers:
compared to my previous best run, the one that resulted in NFNetL1V1-100-0.57141:
And it's all thanks to the folks at TRC, so shout out to them!
I currently have a few runs in progress across a couple of dimensions:
Once the experiments are over, the plan is to select the network definitions that lay on the Pareto curve between throughput and F1 score and release the trained weights.
One last thing.
I'd like to call your attention to the tools/cleanlab_stuff.py script.
It reads two files: one with the binarized labels from the database, the other with the predicted probabilities.
It then uses the cleanlab package to estimate whether if an image in a set could be missing a given label. At the end it stores its conclusions in a json file.
This file could, potentially, be used in some tool to assist human intervention to add the missing tags.
Python
100.0%