Thanks for the follow up. I wish I could afford multiple TB of nvmes but that is unfortunately out of my budget, but it would definitely be better for latency, notice and power draw. This time I will have to stick to HDDs, but I’ll keep looking :) Enjoy your setup!
In the deep learning community, I know of someone using parquet for the dataset and annotations. It allows you to select which data you want to retrieve from the dataset and stream only those, and nothing else. It is a rather effective method for that if you have many different annotations for different use cases and want to be able to select only the ones you need for your application.