jsulz @jsulz.com · Dec 6

I thought that large datasets would be the most interesting, and they can be - here's blanchon/RESISC45, a dataset of 31k images from Google Earth bucketed into 45 taxonomies with 700 photos per taxonomy huggingface.co/datasets/bla...

0 likes 1 replies

?

Replies

jsulz · Dec 6

More interesting is when the directory/file naming convention lets you see the inequity in the bytes; most apparent with multilingual NLP datasets. Yellow sections for directories/files == more bytes devoted to that language. This is facebook/multilingual_librispeech huggingface.co/datasets/fac...