Data and analytics systems
Business data is often one of an organisation’s most valuable assets. But accumulating it creates no value in itself.
Data becomes useful when it lets someone understand a situation better, take a decision, automate an operation or power a product.
The same holds for AI: a strong model does not make up for data that is poorly understood, hard to reach, or whose meaning nobody can quite pin down.
So our approach to data always starts from the same question: what does the organisation need to be able to understand, decide or do? The data system comes after.
Start from the use, not from the data you happen to have
A data project easily starts as an inventory: which databases do we have, which exports can we get, which systems can we connect?
That is not necessarily the right starting point.
We prefer to start from what the system has to make possible: tracking an activity, comparing scenarios, detecting an anomaly, feeding an internal tool, automating a decision or giving context to an AI model.
From that use, you can work back to the data it needs, its sources, the quality it has to reach and the way it will have to be served.
Among other things, it avoids spending months centralising information nobody yet knows what to do with. Centralising all the data is not a data strategy: a useful platform starts from the uses it has to make possible.
A data chain does not stop at the pipeline
Collecting data and moving it from one system to another is only part of the problem.
For it to become genuinely usable, several layers have to work together.
Collect
Data rarely comes from a single place.
Internal applications, business software, partner APIs, files, legacy databases, product events or external sources sometimes have to live together in one system.
The first problem is retrieving that information reliably and repeatably enough.
Structure
The same business reality can be represented very differently depending on which system produced it.
So you have to define how objects relate to one another, which history has to be kept, which transformations are needed, and where the source of truth sits.
Harden
Data that is missing, duplicated, stale or inconsistent can be more dangerous than data that is simply unavailable.
Quality therefore has to be something you can check: freshness, completeness, consistency, provenance and access rights are as much part of the system as the transformations themselves.
Transform
Some data has to be aggregated, enriched, reconciled or computed before it becomes usable.
Those treatments can be deterministic, statistical, or built on machine learning and AI models, depending on the need.
Serve
Data is only useful once it reaches the right place.
Dashboard, internal application, API, automation, AI model or partner export: how data should be served follows directly from what someone has to do with it.
Governance · Security · Observability · Cost
Sources
Applications, APIs, files, events
Collection
Structuring
Quality and semantics
Transformation
Delivery
Analytics
Product
AI
Data has to mean the same thing to everyone
The hardest problems are not always technical.
Two teams can hold exactly the same data and produce two different figures, because they do not put the same reality behind a word like “customer”, “revenue”, “active product” or “conversion”.
So a data platform also has to establish an explicit enough language around the information: what the definition of an indicator is, which source is authoritative, how a figure was computed, which raw data it came from, and who is allowed to see it.
Without that semantic layer, an organisation can have excellent infrastructure and still be arguing about whether its own numbers are right.
Lineage matters for the same reason: when something looks wrong, you have to be able to walk back from the figure on screen to its source and understand the transformations applied along the way.
CRM
ERP
Product
Billing
Active customer
- Definition
- Source of truth
- Calculation rules
- History
Design for consumption, not only for storage
A data system is often thought about from its infrastructure: warehouse, lake, pipelines, transformations.
Its architecture also has to be thought about from the other end: who is going to consume this information, and how?
An analyst exploring data freely, a salesperson who has to decide in thirty seconds, and an AI agent that needs precise context do not have the same needs.
That directly determines the level of aggregation, the freshness required, the acceptable latency, the access rights, the APIs to expose and the way the information is prepared and presented.
It is one of the reasons we treat analytics systems as products in their own right, not merely as infrastructure. The interface counts as much as the pipeline when it is the point of contact between the data and the business.
AI makes data quality matter more, not less
Recent progress in language models can give the impression that any information can now be queried without having to be structured first.
Mostly, it moves the problem.
A model can make complex information easier to reach, bring sources together or produce a summary. But it still has to know which data to use, which of it is reliable, which of it the user is entitled to, and in what context to interpret it.
The more an AI system becomes able to act on company data, the more those questions matter.
That is the direct link with our applied AI and agents expertise: the model brings a new capability for interaction or reasoning, but the quality of the system depends largely on the business data feeding it.
A data platform has to live with the organisation
A pipeline that works today is not finished.
Sources change. Fields disappear. Volumes grow. Business definitions move. New uses appear.
A data system therefore has to make it quick to see that a source has stopped updating, that a transformation is producing an unusual result, that a volume has shifted abruptly, that a job has become too slow or too expensive, or that a change has moved an important indicator.
Observability is not there to produce one more dashboard. It is what tells you whether the data the organisation relies on still deserves its trust.
This is also where the subject meets our product architecture and engineering expertise: automation, reproducible environments, deployment, security, monitoring and documentation are as necessary to data systems as to the applications that consume them.
Build only what has to be built
As with software architecture, we do not believe in a universal data stack.
An organisation does not systematically need a data lake, a real-time platform or dozens of independent pipelines.
The right architecture depends on how often the data changes, the volume handled, the number of sources, the expected uses and the consequences of an incorrect figure.
Some needs are perfectly served by a relatively simple system. Others call for a far more structured chain. The job of engineering is to tell the two apart.
The sophistication of a data platform has to be proportionate to the uses it serves.
What we have already built
These problems come back across several Vezero projects, in very different shapes. The contexts change, but the question stays the same: how do you turn the data available into a system reliable enough for the business to use?
Energy data
For Dsflow, we worked among other things on rebuilding a pipeline processing the energy consumption data of companies.
Analytics tools for the business
Projects such as IQVIA had us working on the visualisation and use of complex business data in healthcare.
Data and specialised models
With Predicity in cosmetics and Foodlytics in food, we work on systems that combine business data with specialised language models.
Data and AI strategy
We also work, alongside sector partners, on engagements further upstream: identifying use cases, assessing what already exists, and building data and AI roadmaps inside large industrial and pharmaceutical organisations.
From data to use
In the end we judge a data system less by how much information it stores than by how usable that information is.
Do users find the right information? Do they understand what it means? Can they trust it? Can it properly feed an application, a decision or a model? And when something changes, is the team able to understand why?
A good data platform does not try to centralise everything. It makes the right data available, understandable and reliable at the moment it becomes useful.
How we frame these subjects, make the trade-offs and commit to them is set out in detail in the Vezero Method.