augmenting large datasets to have more edge case data for training [D]
I have this idea I'd like feedback on. most camera footage for training, is sunny daytime, because that's what cameras record most of the time. The edge cases models actually struggles with, like night, fog, rain or glare, are rare in the data - the idea is to augment it to have more of the rare cases for training
So take a big labeled dataset A and adapt it to look like target B, where the model will actually run. Physics-based effects where possible (fog, rain, low-light noise), a constrained generative model for what physics can't handle (dusk lighting, headlight glare, wet roads), then matching B's camera quality. Labels stay intact throughout.
clear daytime HD driving footage → a cheap dashcam at night in the rain, with glare and heavy compression.
[link] [comments]
Want to read more?
Check out the full article on the original site