Can We Predict the Future Through Data?

In this blog post, we’ll explore how the vast amounts of data we generate in our daily lives are analyzed and utilized, and how past data can help us predict future trends and inform decision-making.

 

How Does Our Daily Life Become Data?

How much data do you think you encounter throughout the day? Thanks to advances in information and communication technology, everything from the KakaoTalk conversations you exchange with friends on your smartphone first thing in the morning, to the little details of daily life and useful information you encounter on social media, to the bus schedules you check for your commute and the use and recharging of your transit card—all aspects of your daily life can become data. We now live in a world where not only these digital activities but also payments made at convenience stores, songs you listen to on your MP3 player, and even the lessons taught by your teachers can be recorded and utilized as data. So, where can this data be used? And we need to consider what benefits it actually offers us.
You’ve probably heard the term “Big Data” at least once through various channels. While a literal translation would be “large data,” Big Data doesn’t simply refer to a large volume of data. Traditionally, Big Data has been explained using the “3Vs”: Volume, which refers to the scale of the data; Velocity, which refers to the speed at which data is generated and processed; and Variety, which means that it includes not only structured data such as numbers and tables but also various forms of data such as documents, images, and videos. While today’s explanations often add factors like data reliability and value, the “3Vs” framework remains widely used as the fundamental way to understand Big Data. Various data analysis techniques, such as data mining, are utilized to identify useful patterns and relationships within this big data and extract meaningful information. The reason the term “big data” has recently emerged as a major issue is closely related to technological advancements. Of course, simple data existed, was recorded, and was managed in the past as well, but the technology and infrastructure required to store and analyze such explosively growing volumes of data on a large scale were insufficient. However, with the advancement of big data technology, methods for storing, processing, and analyzing vast amounts of data are also improving day by day.

 

Can we predict the future using past data?

If so, where can this big data be applied? As the saying goes, “You must understand the past to understand the present and predict the future,” big data analysis can aid in decision-making and be used to predict future outcomes. To give a simple example, by collecting and analyzing data on your eating habits and exercise frequency to date, we can use it to predict various future possibilities—such as what lifestyle patterns you are likely to follow, what foods you are most likely to purchase, and your risk of developing specific diseases. However, these predictions are probabilistic, based on past data and analysis results, and do not definitively determine an individual’s future.
A brief but interesting real-world example of this type of data analysis is wine. Wine is produced by harvesting grapes, undergoing various processes, and aging for a certain period, after which its quality for that year is evaluated by world-renowned wine critics. Ultimately, the quality of wine is often not determined immediately after harvest but is judged after a certain period through expert evaluation. However, by using regression analysis—a representative method of data analysis—it is possible to predict values related to actual expert evaluations in advance. Regression analysis is an analytical method that estimates results by inputting various factors related to a specific value or state in order to predict that value or state. In the case of wine, the value being predicted is the quality. A key aspect of this regression analysis is identifying which input factors influence the outcome and, if they do, to what extent. To determine this, American economist Orley Ashenfelter analyzed various data sets and identified factors related to wine quality in a given year, including winter precipitation from the previous year, the average temperature during the growing season of that year, and precipitation during the harvest period.
Of course, many other factors also play a role, but he conducted a statistical analysis showing that these variables are closely correlated with wine quality, and as a result, he proposed the following equation:

Wine quality for the current year = 12.145 + 0.00117 × winter precipitation of the previous year + 0.06140 × average temperature during the growing season of the current year – 0.00386 × precipitation during the harvest season

This equation made it possible to predict the quality of a vintage based on meteorological data before experts conducted direct wine evaluations, and this analysis became known as a representative example of regression analysis used to predict wine quality and price. Thus, one of the key objectives of regression analysis is to identify the relationship between relevant input variables in order to determine the target value. So, what is the significance of using the above equation to determine quality? Simply predicting quality may not be meaningful on its own, but since wine prices vary significantly depending on quality, and information about a wine’s quality and future value can be relevant to wine trading and investment, it holds significant importance. Therefore, predicting wine quality through such data analysis can serve as a tool for making decisions regarding wine-related transactions or investments. Aschenfelder’s research also demonstrated that data can be used to analyze not only wine quality but also prices and market trends.

 

How can social media data be utilized in corporate product development?

Another example is the use of social media analysis—such as Facebook and Twitter—for corporate new product development. Today, countless posts and reactions appear on social media in real time. Companies can collect and analyze this data over periods ranging from a few hours to several months to identify consumer complaints and design preferences, which they can then use as a reference for developing new products. For example, suppose an analysis of data related to Samsung’s “Galaxy S2” on Facebook and Twitter revealed that many people disliked the sharp corners, complained about excessive heat generation, and noted that the battery drained too quickly. In that case, when developing the next version of the device, the company could utilize this data by focusing on rounding the corners, reducing heat generation, and developing technology to extend battery life. The key takeaway from this example is not simply collecting large amounts of data, but analyzing consumer reactions to derive insights that can inform real-world decision-making, such as product development.

 

Can the future be determined by data alone?

This approach to analyzing big data is already being utilized across various fields, and related technologies are advancing rapidly. Back in 2014, big data-related technologies garnered significant attention, even being highlighted as a key area in the strategic technology trends identified by the global research firm Gartner. Today, however, big data is no longer merely a passing trend; rather, it has established itself as the foundation for data analysis across various industries and societal sectors, integrated with technologies such as artificial intelligence, machine learning, and cloud computing. Therefore, rather than viewing big data as merely a temporary technological fad, it is necessary to understand it as a crucial method for collecting, storing, and analyzing data to inform decision-making.
However, when making decisions based on these data results, one must always keep in mind that data alone should not be trusted 100%. Data can vary depending on the environment and timing of collection, and results can also differ based on the variables used in the analysis or the quality of the data.
Furthermore, since society and human behavior are constantly changing, the results of a single analysis cannot be relied upon indefinitely. Therefore, it is important to clearly recognize that the results of data analysis have certain limitations and to make decisions by comprehensively evaluating the findings alongside other information. Ultimately, the true value of big data lies not simply in collecting large amounts of data, but in making better judgments by analyzing past data to understand the present and predict future possibilities.

 

About the author

Cam Tien

I love things that are gentle and cute. I love dogs, cats, and flowers because they make me happy. I also enjoy eating and traveling to discover new things. Besides that, I like to lie back, take in the scenery, and relax to enjoy life.