Services / 02
Data Engineering
We treat security telemetry as an engineered data estate — sources, pipelines, schemas, lineage, retention and cost — rather than as whatever happens to reach the SIEM. The same discipline applies to the analytical and AI workloads that increasingly consume the same data.
The problem
Security data estates are rarely designed. They accrete. A source is onboarded for one detection, a second team adds a parallel feed, a third builds a report on a field whose meaning changed two years ago. Nothing is wrong in isolation; collectively the estate becomes something no one can reason about — and every downstream capability, detection and AI alike, inherits that.
The symptoms are familiar: an analytic that silently stops matching after an upstream format change; a cost line that grows faster than coverage; a field that means three different things depending on which pipeline populated it; a compliance question that takes two weeks to answer because lineage exists only in people's heads.
Capabilities
What this pillar covers.
Grouped by the part of the system they act on.
Architecture
- Data architecture
- Security data engineering
- Data platforms
- Telemetry architecture
Movement
- Data pipelines
- Data integration
- Normalisation & enrichment
- Routing & tiering
Governance
- Data governance
- Metadata
- Lineage
- Data quality
- Retention & classification
Consumption
- Analytics
- Observability
- Feature and context stores
Approach
How an engagement runs.
The sequence matters more than the tooling. Skipping a stage moves its cost later, it does not remove it.
- 01
Inventory
Catalogue what is actually flowing, including the feeds nobody owns any more. Most estates contain both undetected gaps and expensive duplication.
- 02
Model
Design the schema, the classification and the routing rules. Decide deliberately what is kept hot, what is tiered, and what should never have been collected.
- 03
Build
Implement the pipelines with tests, quality gates and lineage emitted as a first-class output rather than reconstructed later from documentation.
- 04
Govern
Put the operating model in place: ownership, change process, quality SLOs and the metadata that makes the estate navigable by people who did not build it.
What you receive
Concrete artefacts, in your repositories and your platforms.
- A source inventory with owner, schema, volume, fidelity and retention for every feed
- Pipeline architecture: collection, normalisation, enrichment, routing and tiered storage
- A schema and naming standard that survives contact with a second team
- Lineage from source to consumer, so an analytic's dependencies are inspectable
- Data quality checks that run continuously and fail loudly, not quarterly and quietly
- A cost model — which data is expensive, which is load-bearing, and where those two differ
Expected outcomes
Properties of the resulting system — not performance figures, which depend on your estate rather than on us.
- One inventory of sources, with owners, schemas and known fidelity limits
- Pipelines with tests and quality gates, so upstream changes fail visibly
- Lineage that answers 'what breaks if this field changes' in minutes
- A deliberate cost position rather than an accidental one
- Data that AI and analytics workloads can consume without re-deriving its meaning