DSM Distribution’s Michael Jackson in conversation with Lightning IQ’s Nick Pollard

DSM Distribution CRO Michael Jackson talks to Lightning IQ Managing Director Nick Pollard about the hidden risks inside enterprise data estates and why understanding your data is becoming critical to cyber security, compliance, migration and AI
Voices

17 August 2026

In association with DSM Distribution

Most organisations can tell you how much data they hold. But that is very different from knowing what is
contained within that data. As organisations accumulate tens of millions of files across hundreds of terabytes and increasingly petabytes the gap between understanding data volume and understanding data content is becoming a significant business issue.

It is also becoming more urgent as businesses prepare their data for AI. DSM Distribution CRO Michael Jackson (MJ) sat down with Nick Pollard (NP), managing director of Lightning IQ, to discuss why organisations need greater visibility into their data estates, what a data audit can reveal, and why getting the data right may be one of the most important first steps in any enterprise AI strategy.

MJ: Most organisations would probably say they already have a reasonable understanding of their data. What are they missing?

NP: They often understand their infrastructure very well, but that isn’t necessarily the same as
understanding their data.

An organisation might know it has 500TB of storage. It knows which servers or cloud platforms that data sits on and may have sophisticated backup, security and monitoring systems around it.

But ask some different questions.

How much of that 500TB is duplicated? How much hasn’t been accessed for years? How much contains personal or sensitive information? How much has no clear owner? How many different versions of the same document exist? How much is being retained without a valid business reason?

Those questions are much harder to answer. Storage tools tell you about capacity. Backup tools tell you what has been copied. Security tools tell you who has accessed something.

What they generally don’t tell you is what is actually inside the files. That’s the visibility gap we’re trying to address.

MJ: So, is the starting point essentially a data audit?

NP: Exactly, although an important distinction is that we don’t start by assuming what the problem is.
Sometimes a customer approaches the exercise because they think they have a storage problem. Another might be concerned about compliance, cyber security, a migration project or AI.

Lightning IQ connects to live storage across environments including NFS, SMB and S3 and assesses the estate in place. We establish what exists, where it exists, its age, the type of content involved and how the data estate is structured.

From there you can start uncovering things such as duplicate and ROT (redundant, obsolete and trivial) data, dormant files, sensitive information, retention exceptions, unusual concentrations of files, deep folder structures, excessively long file paths and gaps in ownership. The important thing is that the data tells us where the problem is.

A customer may start the project worried about storage growth and discover that its bigger exposure is ageing personal data. Another might expect to find lots of duplicate files but instead discover folder structures that could cause serious problems during a migration or recovery exercise. The audit doesn’t assume the problem. It reveals it.

MJ: Why does scanning the data in place make such a difference?

If you’ve got a relatively small amount of data, moving it around for analysis might be manageable.
Once you’re dealing with tens of millions of files and hundreds of terabytes or several petabytes it becomes a very different proposition.

The traditional approach can involve copying data into another environment, extracting it, expanding containers, indexing it and potentially creating additional working copies before you can even start making decisions about it.

That has an infrastructure cost, but there’s also a security consideration. Every additional copy potentially
increases the exposure surface. Lightning IQ changes the sequence.

Rather than saying, “Move everything and then we’ll work out what’s in it,” we’re saying, “Let’s understand the live estate first and then decide what actually needs to happen.”

I think Scale is one thing, the other is TIME and sheer volume. You might achieve a scan of 100TBs of data but its taken 4 weeks – meaning you’re already out of date (by 4 weeks). Current scanning technology was not built for the Data world we now live in. We’ve been in situations where 100TBs took five days for an incumbent to recognise duplicates. For Lightning IQ it takes a matter of hours.

MJ: And presumably sampling the data isn’t necessarily enough?

NP: That’s right. Sampling can give you an indication, but ultimately it is still an estimate.

If an organisation has 40 million files and you inspect a tiny percentage of them, you can’t necessarily assume the remaining millions have exactly the same characteristics.

When the decisions relate to compliance, sensitive information, cyber risk or a major migration, organisations increasingly want evidence rather than assumptions.

The value comes from being able to build a much more detailed inventory of the live estate and then make decisions based on what is actually there.

MJ: AI must be making this issue significantly more important.

NP: Absolutely. We’re seeing organisations identify enormous document estates that they potentially want to use for AI.

Someone says, “We’ve got seven million documents that our AI could use,” and initially that sounds fantastic. Seven million documents sounds like an enormous corporate knowledge asset. But then you need to ask: what are those 7 million documents?

How many are duplicates? How many are 10 years old? How many are abandoned drafts? How many contain personal or confidential information? How many have been superseded? Which versions were actually approved? Who owns them?

Suddenly your 7 million-document knowledge base looks very different.

The teeth of this is EU AI Regulation. How can you possibly stay on the right side of regulation when you have no clue what is inside those 7 million docs?

MJ: Can you put the scale of that into context?

NP: Seven million documents averaging just 500KB is approximately 3.5TB of source data.

If those documents average 2MB, you’re at roughly 14TB. At 5MB, you’re talking about around 35TB. And real corporate data estates aren’t made up exclusively of neat Word documents.

You’ve got PDFs, scanned documents, presentations, images, e-mail containers, attachments, technical files and all sorts of other formats.

Under a traditional approach, organisations may start copying and processing all of that before they’ve
established whether they actually need it.

You could be paying to move it, extract it, OCR it, index it, store working copies of it, embed it and have people review it. Then you discover later that a substantial proportion should never have been included. That’s an expensive way to find out.

MJ: So, the principle is effectively “reduce before you process”?

NP: That’s exactly it. If you begin with 7 million documents, perhaps analysis determines that 3 million are genuinely relevant. Further governance and approval might reduce that to 1 million trusted sources.

I’m not suggesting those will be the ratios for every organisation, every data estate is different, but the principle is important. Why pay to process millions of files that should never enter your AI workflow?

Understand the source estate first. Remove the noise and unnecessary risk. Then invest the expensive
processing resources in the material that has value.

MJ: There is a lot of discussion around AI hallucination and trust. Is data quality part of that conversation too?

NP: Very much so. People understandably focus on the model, but an AI system can only work with what you give it.

Imagine an organisation has six versions of a policy document. One is current, two are old versions. Another is a draft. Another contains comments that were never approved. If they’re all treated as equally authoritative sources, you’ve created uncertainty before the AI has even answered its first question.

The same applies to duplicated content, inconsistent information, poor metadata and documents with no
identifiable owner or provenance.

For enterprise AI, one of the most important questions is becoming: “Can I trust the source material behind this answer?”

AI readiness therefore isn’t simply about getting documents into a vector database or another AI platform. It’s about understanding which information is current, relevant, approved and appropriate to use.

MJ: Does Lightning IQ prepare or train the AI model?

NP: No, and that’s an important distinction. Lightning IQ doesn’t train the model. Its role is further upstream. We help organisations understand and assess the source estate so they can determine what should move forward into the AI workflow.

Once the appropriate corpus has been identified that selected data can go through whatever controlled process the organisation chooses, extraction, chunking, embedding, vector indexing, retrieval, fine-tuning or another AI architecture.

Our objective is to make sure the organisation isn’t asking the AI to learn from everything simply because
everything happens to be available.

MJ: Beyond AI, where else can the findings make a difference?

NP: There are several areas. Compliance and privacy are obvious ones because businesses may discover personal or sensitive information sitting in unexpected places or being retained beyond the period they anticipated.

Storage optimisation is another. There can be a meaningful financial difference between knowing you have a large storage estate and understanding how much of that estate is providing genuine business value.

Then you’ve got cyber resilience, records management and migration.

During a migration, for example, organisations often don’t want to discover halfway through the programme that they have extreme folder depths, problematic path lengths, huge volumes of redundant content or files nobody can take ownership of.

The output from an audit is therefore not simply an inventory. It’s about giving technical and executive teams a structured view of what was found, why it matters and what they should do next.

MJ: Beyond AI, where else can the findings make a difference?

There are several areas. Compliance and privacy are obvious ones because businesses may discover personal or sensitive information sitting in unexpected places or being retained beyond the period they anticipated.

Storage optimisation is another. There can be a meaningful financial difference between knowing you have a large storage estate and understanding how much of that estate is providing genuine business value.

Then you’ve got cyber resilience, records management and migration.

During a migration, for example, organisations often don’t want to discover halfway through the programme that they have extreme folder depths, problematic path lengths, huge volumes of redundant content or files nobody can take ownership of.

The output from an audit is therefore not simply an inventory. It’s about giving technical and executive teams a structured view of what was found, why it matters and what they should do next.

MJ: If you had to describe the shift you’re trying to create for customers in one sentence, what would it be?

NP: It’s moving the organisation from saying: “We know how much data we have” to being able to say: “We know what our data is, where the risk sits and what action we need to take.”

And for AI specifically, it’s moving from: “We’ve got seven million documents available for AI” to “We know which of those documents are current, relevant, approved and safe to use.”

That distinction is going to become increasingly important.

For years, organisations have invested heavily in storing, protecting and backing up information. The next
challenge is understanding it. As corporate data estates continue to grow and as businesses seek to unlock those estates through AI, visibility into the content itself is becoming increasingly valuable.

For businesses preparing for AI, migration, compliance initiatives or wider data transformation, that may be the more important place to start.


Back to Top ↑