•1 min read•from Machine Learning
[D] How do you get preprocessed dataset of a paper [D]
Hi all,
I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.
I've tried all reasonable interpretations of the filtering described in the paper and the closest I can get is still an order of magnitude off for one of the datasets. The paper says "data available on request" — I emailed the authors and followed up once, no reply so far.
For those who've been in this spot:
- Do you just keep the larger-but-valid version you can reproduce and document the mismatch?
- Is it worth sampling to match the reported size or does that just create a different irreproducible dataset?
- When do you escalate to the journal vs just waiting?
How have you successfully gotten preprocessed files from authors? Any etiquette around follow-ups or journal contacts that actually worked?
Thanks for any advice.
[link] [comments]
Want to read more?
Check out the full article on the original site
Tagged with
#dataset
#preprocessing
#reproduce
#data availability
#raw data
#filtering
#data statistics
#table 1
#sampling
#irreproducible dataset
#authors
#journal
#follow-up
#etiquette
#data request
#mismatch
#machine learning
#reproducibility
#validation
#escalation