Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sure, but publish the numbers so we can decide for ourselves.


They have numbers in the paper[0] linked in TFA. They trained on 29 samples and tested on 12 samples. On those samples, all techniques (NN, SVM, CNN) all had 100% true positive rates. The CNNs did best on false positives, giving on average 2-3 per sample.

Also, it is very important to note: they are comparing a CNN analysing a GC-MS[1] sample to a human expert analysing a GC-MS sample, which is basically a 1D plot of intensities of chemicals. This is not a comparison to a human smelling someone's breath.

Also, as far as I can see, they never compare to a GC-MS run with selective ion monitoring (SIM), which would be a lot faster for the human to analyse.

[0] (ResearchGate link, sorry): https://www.researchgate.net/publication/324921031_Convoluti...

[1] A gas-chromatograph mass spectrometer is a device which samples a gas and detects what chemical species are present. It's what they do at airports when they swipe that piece of fabric on your luggage and put it in a machine (GC-MS) to look for explosives. It's a nice tool, used extensively in labs worldwide, but expensive (usually a couple hundred thousand USD, depending on specifications).


Ah sorry, I was relying on the parent comment's statement that they hadn't published numbers. Hadn't actually checked myself. Glad that they did though.

That being said, those data numbers seem extremely small. In reading the paper, I see that they attempted to augment it a bit. But it still seems near meaningless given the tiny sample set.


Forty-one total samples? Um...its not quite anecdata, but call me when the N is a bit larger.


Maybe obtaining and analyzing these samples is a more complicated process than we think? For example, acquiring and analyzing a malware sample is not equivalent to labeling or classifying an image.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: