1 min readfrom Machine Learning

[D] How do you get preprocessed dataset of a paper [D]

Hi all,

I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.

I've tried all reasonable interpretations of the filtering described in the paper and the closest I can get is still an order of magnitude off for one of the datasets. The paper says "data available on request" — I emailed the authors and followed up once, no reply so far.

For those who've been in this spot:

  • Do you just keep the larger-but-valid version you can reproduce and document the mismatch?
  • Is it worth sampling to match the reported size or does that just create a different irreproducible dataset?
  • When do you escalate to the journal vs just waiting?

How have you successfully gotten preprocessed files from authors? Any etiquette around follow-ups or journal contacts that actually worked?

Thanks for any advice.

submitted by /u/Individual-Safety906
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#dataset
#preprocessing
#reproduce
#data availability
#raw data
#filtering
#data statistics
#table 1
#sampling
#irreproducible dataset
#authors
#journal
#follow-up
#etiquette
#data request
#mismatch
#machine learning
#reproducibility
#validation
#escalation