Public API

Merit Capital Data API

This API contains the full database powering Merit Capital, queryable without authentication.

OpenAPI spec for agents
/data
Machine-readable; regenerated on every schema change.

Example queries

Filtering, ordering, column selection, and pagination follow the PostgREST query syntax, including resource embedding between investors and companies in both directions.

Reference

Generated live from the database schema; the column descriptions are the same ones served to agents at /data.

GET/data/investors

All venture investors, one row each: institutions and individuals share the base profile fields; the institutional fields (aum, funds, individuals, founding_year) and the individual field (affiliation) are populated depending on investor_type, with the non-applicable side null.

Identity

id
text
primary keyStable unique identifier (text).
name
text
Display name of the investor.
slug
text
URL-friendly unique identifier. Use for lookups, e.g. investors?slug=eq.sequoia-capital.
investor_type
enum
"Institutional" or "Individual". Determines whether aum/funds/individuals/founding_year (institutional) or affiliation (individual) are populated.
icon_url
text
URL of the investor's logo/icon image.
website
text
Official website URL.
hq_location
text
Headquarters location as free text.

Investment thesis

investment_concept
text
Investment philosophy / thesis description.
history
text
Narrative history of the investor.
investment_stage
json
JSON array of stages, e.g. ["Seed","Series A"].
investment_geo
json
JSON {core: string[], other: string[]} geographic focus.
investment_scope
json
JSON {type: "Generalist"|"Specialist", focus: string[]}.
investment_range
json
Check size USD JSON {lower_limit, upper_limit, median}; null where the investor's check size is not known.
notable_investments
json
JSON array of notable portfolio company names.

Performance

track_record
json
derived from companiesread-onlyJSON {total_companies_invested, investing_start_year, one_b_companies | ten_b_companies | hundred_b_companies | ipo_companies: {amount, names[]}}. Cumulative: a $50B public company appears in one_b, ten_b and ipo. ipo_companies counts companies currently public (status = public), not ever-IPOed. investing_start_year is the investor's founding_year, not derived from companies. total_companies_invested counts companies carrying a determined outcome; rows whose company could not be identified are excluded.
graduation_rate
json
derived from companiesread-onlyJSON {seed_to_series_a: {rate_percent, graduated, sample_size, seed_entries, censored}}; null when the investor has no Seed-stage entry. seed_entries counts every Seed-stage entry, each of them decided; censored are too young to count (entered under 24 months ago); sample_size is the rest, the cohort, every one of them graduated or not_graduated; rate_percent is graduated / sample_size, null when the cohort is empty. Coverage of the cohort is complete by construction: a Seed entry past the 24-month clock with no evidence of a later round is counted as not graduated rather than left out. The rows behind the counts: /companies?investor_id=eq.<id>&entry_stage=eq.Seed.

Portfolio distributions

company_stage_distribution
json
derived from companiesread-onlyJSON array [{stage, companies}] sorted by companies desc, ties alphabetical; sums to the investor's companies row count. Stage vocabulary is companies.entry_stage. Share is computed by consumers, never stored. The rows behind the counts: /companies?investor_id=eq.<id>.
company_geo_distribution
json
derived from companiesread-onlyJSON {country: [{country, companies}], city: [{city, companies}]}; each sub-array sums to the investor's companies row count and sorts by companies desc, ties alphabetical, null element last. A null-keyed element is a real count: null city means the HQ is unresolved. Countries are ISO 3166-1 alpha-2; cities are bare municipalities. Share is computed by consumers, never stored. The rows behind the counts: /companies?investor_id=eq.<id>.
company_focus_distribution
json
derived from companiesread-onlyJSON array [{focus, companies}] over the closed twelve-value focus taxonomy, sorted by companies desc, ties alphabetical, null element last; sums to the investor's companies row count. The null element counts companies outside the taxonomy and is a real published number. Share is computed by consumers, never stored. The rows behind the counts: /companies?investor_id=eq.<id>.

Institutional detail

aum
json
institutions onlyUSD JSON range {bottom, top}. Null for individuals or when undisclosed.
funds
json
institutions onlyJSON wrapper object {funds: [{name, size (USD number), vintage (year)}]}. The array sits under the "funds" key, not at the top level. Null for individuals or when undisclosed.
individuals
json
institutions onlyJSON array of the people at the firm, elements shaped {name, position, socials: {linkedin_url, x_url}}; position and either social URL may be null. Null for individuals.
founding_year
number
institutions onlyyear the firm was founded.

Individual detail

affiliation
json
individuals onlyJSON {current: [...], past: [...]} affiliations with roles and years. Null for institutions.

Metadata

created_at
timestamp
Row creation timestamp.
updated_at
timestamp
Last write timestamp; doubles as last-verified under the verify-then-ingest playbook.

GET/data/companies

Portfolio companies as classified by the outcome pipeline, one row per (investor, company) pair. A company backed by two investors appears twice. Also the derivation source for the investor-level metrics: recompute_investor_metrics aggregates these rows into investors.track_record, graduation_rate and the three company_*_distribution columns on every write to this table. Filter one portfolio by investor_id, or by slug through an embedded filter: ?select=*,investors!inner(slug)&investors.slug=eq.<slug>. Embedding investors<->companies works in both directions, but embedded lists carry no Content-Range; paginate on this endpoint, never inside an embed.

Identity

id
uuid
primary keySurrogate key, minted by the database on insert. Ingest never sends it.
investor_id
text
The investor whose portfolio this row belongs to; references investors.id.
name
text
Company name as the investor's portfolio renders it. Unique per investor; the ingest conflict key.

Classification

hq_city
text
Headquarters city, free text, bare municipality ("Palo Alto", never "Palo Alto, CA"). Null means unresolved, not "no HQ".
hq_country
text
Headquarters country, ISO 3166-1 alpha-2 code (US, IL, ...).
focus
text
Sector focus from the closed twelve-value taxonomy. Null means the company falls outside the taxonomy, not "unknown".
entry_stage
text
The stage label of the round in which the investor first put fund money in: Pre-Seed, Seed, Series A, Growth, or `undetermined` where the record does not give the round. `undetermined` excludes the row from the stage distribution and from the seed-to-A cohort, and from nothing else — an unknown ROUND is not an unknown COMPANY, so the row keeps its band, its status, its country and its place in the track record. A row that carries `undetermined` because the company itself could not be identified reads `undetermined` in valuation_range too, and that is what takes it out of every metric.

Outcome

status
text
Current company state: private, public, acquired, defunct or unknown.
valuation_range
text
Cumulative valuation bucket: below_1b, 1b_plus, 10b_plus, 100b_plus. `undetermined` means the company could not be valued, which is always true when it could not be identified and can also be true on its own: a Form D sequence or a programme's standard terms give an entry stage for a company nobody can name, and such a row carries a stage with no band. A row with any undetermined column carries no valuation_usd and is excluded from every derived metric. A bucket is a judgment, unlike the measured valuation_usd.
valuation_usd
number
Exact current valuation in USD, source-gated. Null means no sourced figure, not zero.
valuation_asof
date
Date the valuation_usd figure was current; always read the two together.
seed_to_a_graduation
public.graduation_status
Seed-to-Series-A graduation of a Seed-stage entry, as a status, and every Seed entry carries one. graduated: a later priced round, cited or present in the company's own Form D sequence; a later round named Series A or later counts whatever its size. not_graduated: no later round, established by a dated search of named surfaces (the firm's own archive, the company's Form D sequence where it files in the US, and at least two publisher surfaces) at least 24 months after entry, by Form D absence, or by a cited acquisition or shutdown. censored: entered less than 24 months ago, too young to count. undetermined appears only on a row whose entry_stage is undetermined too: a company that could not be identified, or a token-funded position with no priced round to graduate from; neither is a Seed entry. Null wherever else the column does not apply, a non-Seed entry. Every graduated or not_graduated row is verified (a citation, the filings, or the recorded search) before it is ingested.

Metadata

created_at
timestamp
Row creation time.
updated_at
timestamp
Last ingest touch; a re-run that changes nothing else still bumps this.