Big Data: Humble Objectives

A Probability Reminder

Many discussions about Big Data begin with the assumption that complex phenomena require complex tools. Vast storage systems, distributed computing clusters, machine learning models, and real-time analytics engines are frequently presented as the natural response to modern business complexity. Yet some of the most surprising insights about uncertainty and variability emerge from extremely simple probability problems that can be solved using nothing more than high-school mathematics.

Consider the well-known Birthday Paradox (technically called the Birthday Problem, though “paradox” sounds more dramatic and therefore receives better marketing in textbooks). The result is famously counterintuitive. With only twenty-three randomly chosen people in a room, there is already about a fifty percent probability that two individuals share the same birthday, ignoring the year of birth. Increase the group to seventy individuals and the probability rises to more than ninety-nine percent. With three hundred and sixty-six people present, the probability reaches certainty because there are only three hundred and sixty-five possible days.

The mathematics behind this result requires nothing more than basic probability calculations, yet the outcome contradicts most people’s intuition about randomness. In a similar vein, the Gambler’s Ruin Problem (a classical probability model describing why someone repeatedly betting with limited capital almost inevitably loses against a player with larger resources, which is the mathematician’s polite way of explaining why casinos continue building larger hotels every year) demonstrates how seemingly favourable odds eventually collapse under repeated trials.

Neither of these problems requires advanced statistics, supercomputers, or sophisticated analytics platforms. They emerge directly from first principles.

The Big Data Alternative

One can imagine an alternative universe in which humanity had never discovered these probability results. In that world researchers might attempt to prove the Birthday Problem empirically by collecting enormous quantities of data. Thousands of experiments would be conducted. Observations would be recorded across different cities, climates, and demographic groups. The results might be stored in a vast columnar data warehouse containing hundreds of petabytes of observational data. After months of analysis someone might finally conclude that in nine out of ten groups of fifty random people at least one pair shares a birthday, while certain regions show slightly higher frequencies because of elevated twin birth rates.

Such a project would look impressively modern. It would also be completely unnecessary.

The point is not that Big Data is useless. The point is that data volume alone does not create understanding.

A Supply Chain Perspective

My own professional training lies in supply chain management rather than mathematics or statistics. Supply chains are systems defined by variability, uncertainty, and what might politely be described as systemic corruption by randomness (meaning that even perfectly designed processes eventually behave in strange ways once real customers, machines, weather conditions, and transport delays are introduced).

Years of observing operational data tend to sharpen intuition about patterns within that variability. One begins noticing recurring relationships that are not immediately visible in planning models.

For example, a planner may notice that every time a particular product is ordered between December and January, service levels deteriorate because the product requires a specific machine that is already operating near full capacity. Producing that product therefore delays several other items that depend on the same resource.

Another pattern might reveal that customers from a particular region almost always order product A whenever they purchase products B and K, but not when they purchase B and D. A marketing promotion may reveal that discounting product X together with product Y increases gross margins more than discounting either product individually. A planner might observe that a five percent price discount reduces profitability sharply while a twenty percent discount increases profitability slightly because the resulting sales volume compensates for the lower price.

None of these insights require petabytes of data. They arise from careful observation and domain understanding.

Observations Become Strategy

Such observations, although informal, can influence real business decisions. A company might rationalise its product portfolio by eliminating certain SKUs in order to improve delivery performance. Promotional strategies may become more targeted once product substitution patterns become visible. Production schedules may be adjusted to avoid combinations of products that consistently create bottlenecks.

Competitive analysis may reveal that a product competes not only with similar items in its category but with entirely different product types that satisfy the same consumer need. Chocolate cookies may compete with dark chocolate bars if the price difference remains small enough.

Insights such as these transform operational decision-making even before advanced analytics enters the discussion.

The Real Meaning of Big Data

After reading numerous articles and presentations on the subject, the most useful definition of Big Data I have encountered is surprisingly modest. Big Data is not primarily about collecting more information. It is about discovering previously unknown relationships within the information already available.

In the language of analytics practitioners this is sometimes described as identifying “unknown unknowns” (a phrase popularised by policymakers but now enthusiastically adopted by data scientists, partly because it sounds profound and partly because no one can argue with something that is, by definition, unknown).

Yet even this objective benefits from discipline.

The Hypothesis Problem

Before embarking on large-scale data initiatives, organisations should ask a simpler set of questions.

Do we already possess hypotheses derived from existing operational data? Have those hypotheses been tested through practical experiments within the business? If an initiative based on those hypotheses failed to produce expected results, was a sensitivity analysis conducted to determine whether the underlying assumptions were incorrect?

These questions impose intellectual discipline on data initiatives.

Without such discipline, Big Data programmes often begin with enormous enthusiasm for data collection but very little clarity about how that data will actually influence business decisions.

Storage Is the Easy Part

Modern technology has largely solved the problem of storing and processing large volumes of data. Distributed computing platforms, cloud infrastructure, and real-time analytics systems make it possible to capture vast streams of structured and unstructured information. Data storage costs continue to decline steadily.

In the near future it is entirely plausible that large data marketplaces will emerge where organisations can purchase aggregated datasets on subscription, much like purchasing electricity from a utility.

The technical challenge of collecting data is therefore no longer the primary constraint.

The Real Challenge

The more difficult question concerns use cases. What meaningful problems should Big Data solve within real business environments?

In supply chains this might include predicting complex substitution patterns between products, identifying early indicators of competitor activity, detecting subtle relationships between promotional campaigns and regional demand shifts, or modelling how disruptions propagate through supplier networks.

In broader business functions similar opportunities may exist in areas such as customer behaviour, pricing strategy, workforce engagement, or operational risk.

Each use case begins not with data collection but with a hypothesis about how the system behaves.

Humble Objectives

Big Data therefore benefits from modest expectations.

Its purpose is not to replace human judgement or domain expertise. Its purpose is to reveal patterns that were previously invisible, patterns that can then be interpreted by people who understand the system in which those patterns occur.

When used in this way, data becomes a tool for discovery rather than a substitute for thinking.

And discovery, unlike storage, still requires curiosity.