What the dataset carries, what is decided, what is still open. The directory itself is on the landing page.
Everything on the directory is editable by your team, without us. gd.sanerebels.com/studio
Frederic's CSV, dissected: 18 columns, 1,000 collaboration rows from the ~400,000-row crawl. Every column measured, every column given a verdict. Counted 2026-08-14.
The full field list of the sample, measured on 1,000 rows. Percentages are how often a column carries a usable value.
elimity.comipool.sesncb.be ica.seWalmart is a retail corporation that operates several chains of discount departments and warehouse stores. ICA is an online shop that offers ICA's lunch box, healthy food and catering services.{e-commerce,"food and beverage","supply chain management"} {e-commerce,grocery,retail,"retail technology",shopping}Elimity helps their customers to protect their important information assets with superior identity governance solutions. Ipool is a user-friendly application that supports workforce management.Elimity eliminates identity-based risks and implements mature identity governance… Eoliann harnesses satellite data and proprietary ML algorithms to predict climate risk…2017 2010Mechelen, BEL Stockholm, SWEBEL SWE1-10 11-5010 19893076 1596572{Tenity,"Gemma Frisius Fund"} {"Primo Capital","Exor Ventures","Compagnia di San Paolo"}{"cyber security","information technology","risk management"} {"facilities support services",software}{"information technology","privacy and security"} {"administrative services",software}[{"type": "pilot", "domain": "workday.com"}] [{"type": "pilot", "domain": "rabobank.com"}]thoropass.com/customers/rimidi/ workato.com/customers/nutanix2023-11-30 06:24:40+01 2025-07-28 08:02:44+02Both domain columns are 100% filled, and the domain is the identity key everywhere. Corporate descriptions 92%, startup descriptions 100%, founded year 100%, startup HQ 99%, country 96%. Plus 61 rows with a public case-study link: the provenance we link instead of asserting collaborations ourselves.
Industries and categories (90% / 95%) are messy crawl tags and reach the UI only through the deterministic taxonomy. Funding (39%) sorts the works-with examples and is never published as a company figure. Employee ranges (98%) beat the numeric counts (45%). Investors (40%) and referenced customers (18%) enrich where they match and stay silent where they do not.
partnership_created_at looks like a date and is a crawl artifact: 40% of rows sit on two days in July. No feature touches it. Partnership types resolve on 5% of rows, too thin to filter or badge on. Long descriptions (73%) add nothing over the short ones.
Who buys most: Amazon 18, Google 17, Microsoft 14. Which industries buy: Retail & Consumer 180, Technology 154, Automotive & Mobility 133, Financial Services 125, Telecom & Media 84. Where the startups sit: USA 319, UK 94, Germany 60, France 58, India 39, Canada 36. How concentrated: barely, the top 10 hold 9% and 528 of 695 corporations show exactly one collaboration.
Built and running on stand-in coordinates already: HQ pins as a second map layer, the countries stat, and the region filter. Frederic's file swaps the coordinates and widens the coverage; still to come are region x industry slice pages ("Top Financial Services, Germany"), buyer country against startup country, and the real cutoff (3 or 5) from the full distribution.
Trends over time: the dates are crawl timestamps. Deal sizes: not collected. Who to call: company and claimed state only. Completeness: the crawl finds what companies published, so every count is a floor.
Settled in the call on 2026-08-12 and in the days after. Each one is live in the product.
The small calls of one working day, so nothing has to be reconstructed from the build.
The constraints the build has to hold, independent of features.
Calls that are not ours to make, or that need the full dataset first.
Explorations from the scoping week. Kept for reference, not maintained.