Dark Reading is part of the Informa Tech Division of Informa PLC

This site is operated by a business or businesses owned by Informa PLC and all copyright resides with them.Informa PLC's registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. Number 8860726.

Threat Intelligence

6/13/2017
06:15 PM
Connect Directly
Twitter
LinkedIn
RSS
E-Mail
50%
50%

How Bad Data Alters Machine Learning Results

Machine learning models tested on single sources of data can prove inaccurate when presented with new sources of information.

The effectiveness of machine learning models may vary between the test phase and their use "in the wild" on actual consumer data.

Many research papers claim high rates of malware detection and false positives with machine learning, and often deep learning, models. However, nearly all of these rates are within the context of a single source of data, which authors use to train and test their models.

Machine learning has become more advanced but isn't used enough yet in security, says Hillary Sanders, data scientist for Sophos' data science research group. She anticipates usage will increase in coming years to address the rise of different forms of malware.

Historically, Sanders explains, static signatures have been used to detect malware. This method doesn't scale well because software needs to be updated with new signatures as more malware is created. Machine learning and deep learning automatically generate more flexible patterns, which could better detect malicious content compared with stricter static signatures.

"This enables us to move away from signature detection and more toward deep learning detection, which doesn't really require signatures and is going to be better at detecting malware that has never been seen before," she says.

The challenge is in creating a deep learning model to detect forms of malware that don't yet exist. Sanders explains the problem of using current data to test these models, which would ideally be used to detect future malware strains in different clients and environments.

"We can't be sure the data we trained on is going to be super similar to the data in organization deployment," she explains. "If we're training on data that isn't like the data we want to eventually test on, our model might fail catastrophically."

In current machine learning research, accuracy estimates don't consider how systems will process future data. Sanders says modern publications lack time-decay analysis and sensitivity analysis, which could lead to a lack of trust among those who rely on this information.

"If researchers forget to focus on sensitivity testing and time decay, our models are liable to fail catastrophically in the wild," she explains.

Time-decay analysis simulates how the accuracy of data decreases over time, she explains. Consider a dataset with information from January through April. If a machine learning model is trained on data before February 1, it will do well on processing data from January, but accuracy will begin to decay after February.

Sensitivity analysis tweaks inputs for machine learning models to see how output is affected. Sanders will present sensitivity results in her presentation titled "Garbage In Garbage Out: How Purportedly Great Machine Learning Models Can Be Screwed Up By Bad Data" at this year's Black Hat USA conference in Las Vegas.

This analysis will include a deep learning model designed to detect malicious URLs, which was trained and tested using three sources of URL data. As part of her discussion, she'll dive into what caused the results by focusing on how the data sources are different, and higher-level feature activations the neural net identified in some datasets but not in others.

For security teams, the end goal with deep learning is to stop malware. If training and testing data is biased compared with real-world data, models are likely to miss out.

"You ignore the thing you could be optimizing for," says Sanders. "You could miss swaths of malware."

Black Hat USA returns to the fabulous Mandalay Bay in Las Vegas, Nevada, July 22-27, 2017. Click for information on the conference schedule and to register.

 

Related Content:

Kelly Sheridan is the Staff Editor at Dark Reading, where she focuses on cybersecurity news and analysis. She is a business technology journalist who previously reported for InformationWeek, where she covered Microsoft, and Insurance & Technology, where she covered financial ... View Full Bio
 

Recommended Reading:

Comment  | 
Print  | 
More Insights
Comments
Newest First  |  Oldest First  |  Threaded View
COVID-19: Latest Security News & Commentary
Dark Reading Staff 10/13/2020
Where Are the 'Great Exits' in the Data Security Market?
Dave Cole, Cofounder and CEO, Open Raven,  10/13/2020
Overcoming the Challenge of Shorter Certificate Lifespans
Mike Cooper, Founder & CEO of Revocent,  10/15/2020
Register for Dark Reading Newsletters
White Papers
Video
Cartoon
Current Issue
Special Report: Computing's New Normal
This special report examines how IT security organizations have adapted to the "new normal" of computing and what the long-term effects will be. Read it and get a unique set of perspectives on issues ranging from new threats & vulnerabilities as a result of remote working to how enterprise security strategy will be affected long term.
Flash Poll
How IT Security Organizations are Attacking the Cybersecurity Problem
How IT Security Organizations are Attacking the Cybersecurity Problem
The COVID-19 pandemic turned the world -- and enterprise computing -- on end. Here's a look at how cybersecurity teams are retrenching their defense strategies, rebuilding their teams, and selecting new technologies to stop the oncoming rise of online attacks.
Twitter Feed
Dark Reading - Bug Report
Bug Report
Enterprise Vulnerabilities
From DHS/US-CERT's National Vulnerability Database
CVE-2020-15256
PUBLISHED: 2020-10-19
A prototype pollution vulnerability has been found in `object-path` <= 0.11.4 affecting the `set()` method. The vulnerability is limited to the `includeInheritedProps` mode (if version >= 0.11.0 is used), which has to be explicitly enabled by creating a new instance of `object-path` and settin...
CVE-2020-15261
PUBLISHED: 2020-10-19
On Windows the Veyon Service before version 4.4.2 contains an unquoted service path vulnerability, allowing locally authenticated users with administrative privileges to run malicious executables with LocalSystem privileges. Since Veyon users (both students and teachers) usually don't have administr...
CVE-2020-6084
PUBLISHED: 2020-10-19
An exploitable denial of service vulnerability exists in the ENIP Request Path Logical Segment functionality of Allen-Bradley Flex IO 1794-AENT/B 4.003. A specially crafted network request can cause a loss of communications with the device resulting in denial-of-service. An attacker can send a malic...
CVE-2020-6085
PUBLISHED: 2020-10-19
An exploitable denial of service vulnerability exists in the ENIP Request Path Logical Segment functionality of Allen-Bradley Flex IO 1794-AENT/B 4.003. A specially crafted network request can cause a loss of communications with the device resulting in denial-of-service. An attacker can send a malic...
CVE-2020-10746
PUBLISHED: 2020-10-19
A flaw was found in Infinispan version 10, where it permits local access to controls via both REST and HotRod APIs. This flaw allows a user authenticated to the local machine to perform all operations on the caches, including the creation, update, deletion, and shutdown of the entire server.