Show, don't tell.
What we've built.

No vague promises, just results.1 Below is a selection of the systems, data pipelines and models we've designed and built.

1 And some retro ASCII art, because we couldn't help ourselves.

01

High-Dimensional Market Clustering

The challenge

A client wanted to find companies that operate like their best-performing locations, not by the sector code on paper, but by how they actually work.

The build

We built a model that compares companies based on their websites and social media, not on their sector label. So you see which companies truly resemble each other.

The outcome

A clustering engine that groups 47,000+ companies by how they actually operate instead of by sector codes. The client now finds sites that resemble their best-performing locations, not their official industry peers.

Doesn't generalize to

Similarity degrades when companies have sparse observable data in the registry. Works best in well-populated national datasets with rich operational signals.

02

Automated Façade Intelligence

The challenge

Assessing thousands of potential retail sites meant expensive, error-prone field visits.

The build

We built a computer-vision model that pulls façade features from street imagery automatically.

The outcome

80% less field work. 86,000+ sites assessed from street imagery, with consistent features instead of scores that varied per surveyor.

Doesn't generalize to

Confidence drops in areas with limited street-level imagery coverage or outdated Google Street View data. Manual verification recommended for sites with thin image history.

03

Precision Retail Risk Modeling

The challenge

Retailers picked new store locations on gut feeling and scattered local data.

The build

We brought hundreds of data sources together into one risk score for every retail property in the Benelux, the same measure for every location.

The outcome

One comparable risk score for every commercial property in the Netherlands, 300,000+ properties built from 1,200+ data sources. The client ranks sites on a single measure instead of gut feeling.

Doesn't generalize to

Calibrated for the Dutch market. Applying profiles to other countries requires retraining on local property and footfall datasets.

04

Predictive Asset Deployment

The challenge

Security was deployed reactively and often arrived too late at high-value sites.

The build

We built a model that reads behavioural and environmental data to spot high-risk sites in advance, so security is there before anything happens.

The outcome

From reactive scouting to scoring up front: the model narrows 12,400 sites to the 412 with the highest odds, lifting conversion by 34%.

Doesn't generalize to

Accuracy drops when there are insufficient historical deployment outcomes for calibration. Major urban development can disrupt the behavioral patterns the model learned on.

05

Multi-Modal Traffic Measurement

The challenge

Media companies priced outdoor advertising on dated, periodic traffic estimates.

The build

We combine smartphone pings with infrared imagery into a live picture of how many people are passing right now.

The outcome

Live footfall per location, refreshed within 2 seconds. The client prices outdoor media on current traffic instead of dated estimates.

Doesn't generalize to

Smartphone pings systematically underrepresent certain demographic groups. Signal is thinner in car-dependent areas or where smartphone penetration is lower.

06

IRIS: Strategic Location Intelligence

The challenge

One platform for location intelligence and revenue forecasts, built to last.

The build

We built IRIS: our own infrastructure that turns large volumes of location data into substantiated revenue forecasts.

The outcome

One platform that turns location data into revenue forecasts, in use across 25+ markets. The infrastructure Blink and our projects run on.

07

Blink: Custom Spatial Embeddings

The challenge

Standard models miss the invisible spatial links (walking routes, neighbourhood dynamics) that really drive success.

The build

Blink is our own deep-learning model that translates European cities into usable signals for portfolio choices and revenue forecasts.

The outcome

One model that scores 1.2 million+ locations across Europe, from QSR revenue to parcel-locker placement.

08

Text Anonymisation

The challenge

Municipalities had millions of sensitive records they couldn't legally use for analysis under the GDPR.

The build

We built a model that automatically recognises and removes personal data from free text, so the data becomes safe to use for research.

The outcome

4.2 million records stripped of personal data, so the municipality can use information that was legally locked before. ~350,000 documents per hour.

Doesn't generalize to

Optimised for Dutch administrative text. Multilingual or highly unstructured documents require additional domain-specific training.

Something like this on your list?
Let's talk.

The Big Data Company B.V.
Princetonlaan 6 · 3584 CB Utrecht, Netherlands
+31 30 899 9477