Access and Feeds

Synthetic Data: Fake Data Can’t Match the Richness of the Real Thing

By Dick Weisinger

Synthetic data is artificially-generated data that is used to train AI and machine-learning algorithms. Rather than measure and collect data from real events, synthetic data is derived from a small number of existing data sets.

Training machine learning algorithms often requires significant amounts of data. Synthetically-generated data can enable small and mid-sized companies that don’t have the vast data of companies like Google and Facebook, to still be able to train machine-learning algorithms. It’s an attractive proposition.

A common way to create synthetic data sets is to use existing data to build statistical models of the underlying process. The new data set is built by selecting data from probability distributions which model the parameters of the process.

But does it really work? It can. For certain applications, particularly those related to imagining and vision, synthetic data can effectively fill the gap of of having only a small data set. But for other applications, synthetic data isn’t a good choice. Machine learning often uncovers unexpected relationships and parameters that exist in a data set. Data created synthetically simply isn’t as rich.

Alexandre Gonfalonieri, AI Strategy consultant, said that “depending on the nature of the project, I believe that if you understand the intended data well enough to generate an essentially perfect synthetic data set, then it becomes pointless to use machine learning since you already can predict the outcomes.”

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*