Author: IRPA AI Analyst & Senior Advisor, Chris Surdak

To say that the Chinese firm DeepSeek made a bit of a splash on the world stage last week may be the understatement of the decade. The claims made in relationship to their new generative AI Large Language Model (LLM) are being scrutinized by many and if proven true are extremely disruptive to the multi-trillion-dollar AI market. These claims include assertions of multiple-orders-of-magnitude reductions in training costs while producing better results by many metrics. Regardless of whether or not all of these claims are true, what is undeniable is that many of the assumptions that have underpinned the hyper-scale AI market are being called into question.

Data Might Be the Key.

As with probably thousands of others right now, I'm deeply-digesting the DeepSeekMath whitepaper. This seemingly-innocuous line really resonated with me: "In the fourth iteration. we noticed that nearly 98% of the data has already been collected in the third iteration, so we decide to cease data collection."

Anyone interested in the guts of how DeepSeek drove down their training costs should review section 2.1 of this whitepaper, titled “Data Collection and Decontamination.” This section outlines a well-structured campaign by DeepSeek to tightly control the data which made it into the training corpus. Rather than trying to Hoover up the entire Internet in a “Bigger must be better” approach, DeepSeek sought to ONLY ingest data that was:

A. Relevant

B. Accurate

C. Informative

D. Contextually-rich

While these efforts represented a significant amount of pre-training effort, they reduced the total quantity of data that had to be consumed in training by orders of magnitude.

Further, DeepSeek reduced the resolution of the data captured and their resulting training weights, in recognition that dozens of additional significant digits in these weights represent little pragmatic value in comparison to the amount of computational overhead they incur.

Regardless of the other controversial aspects of DeepSeek’s claims (such as how they may have obtained large numbers of GPUs, and if so how many), what remains from their documentation is a recognition that front-end data efficiency can lead to dramatic improvements in processing efficiency. This is as true for a seemingly-tiny company like DeepSeek as it is for an AI behemoth like OpenAI. At this point, it may become an inarguable truth that wholly pedestrian techniques like data classification, deduplication and near-deduplication may fundamentally transform the AI era. 

It was widely reported how Elon Musk stated 2 weeks ago that we are basically out of data. Sam Altman said last month that making progress is proving harder despite massive investments. Both AI pioneers have stated that artificially creating vast quantities of synthetic data may be required to further train artificial intelligence; a seemingly circular-firing-squad approach to achieving “artificial intelligence.” 

Instead, perhaps data quality is the key to effective AI rather than data quantity. If someone wishes to learn to speak French, it is unlikely that taking lessons in Kanji will usefully contribute to that goal. In the same vein, tightly and comprehensively controlling the inbound training of generative AI may be the key to both improved results and lower costs; a win-win scenario for everyone.

As many of us were saying almost 2 years ago, data classification may well be the key to achieving Artificial General Intelligence (AGI). Don't try to learn by ingesting *EVERYTHING*, especially if it's coming from the Internet; in aggregate the worst example of human communications that has ever existed. Rather, know what you're ingesting, why you're ingesting it, and know the difference between "Good" and "Bad" before you start. That this approach seems to work dramatically better should not come as a surprise.


About the Author


Chris Surdak is a Senior IRPA AI Advisor and was formerly White House Chief Transformation officer, Automation & AI Practice Lead at EY & Executive Partner for Digital Transformation at Gartner. He’s an engineer, futurist, transformation executive and best-selling author, with over 30 years’ experience in technology development and deployment, digital transformation, blockchain, data and analytics and AI & intelligent automation.
 


CLICK HERE TO SCHEDULE AN ANALYST CHAT

Links:


Originally posted in the IRPA AI Network — Enterprise AI