1 min readfrom Machine Learning

UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]

I’m working on performance regression detection using machine learning/anomaly detection.

My setup is basically:

  • Healthy runs are used to learn normal behaviour
  • Regression runs are used to see whether the model detects the anomaly
  • For each counter group I only have about 10 healthy samples
  • I’m currently using leave-one-out on the healthy data to set the detection threshold
  • The regression samples are not used during training or threshold selection

I’m confused about a few things:

  • Do I still need a normal train/validation/test split for this type of one-class anomaly detection?
  • With only 10 healthy samples, is leave-one-out better than splitting them into something like 60/20/20?
  • Can the regression samples simply act as the unseen test set?
  • Would it be better to collect a second independent healthy dataset and use that as a final test for false positives?
  • For evaluation, should I mainly use false-positive rate and detection rate/recall rather than MSE/MAE, since I’m not predicting a continuous value?

Just trying to make sure the evaluation setup is correct before I finalise it.

submitted by /u/ZeroDark_Hereford
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#performance regression detection
#machine learning
#anomaly detection
#hardware counters
#leave-one-out
#one-class anomaly detection
#false positives
#detection rate
#recall
#train/validation/test split
#threshold selection
#healthy data
#regression samples
#evaluation
#continuous value
#MSE
#MAE
#counter group
#independent dataset
#normal behaviour