The Original Data You Already Have
The most citable thing most organisations own is a number only they can produce. Not a survey they’d have to commission, not a study they’d have to fund — a number that already exists inside a system somebody on the team logs into every day.
Nobody else can publish it, which is the whole point. A fact that only you can report is a fact that has to be attributed to you.
Where original numbers hide
Run this as an inventory, not a brainstorm. Go through the systems you operate and ask what each one knows in aggregate.
- Your product’s usage patterns. What people do with your tool, in what order, how often, and where they stop. Aggregate behaviour across your customer base describes an industry in a way nobody outside your company can see.
- Support ticket categories. What people ask about, ranked. One of the richest and most neglected datasets in any business: a direct measurement of what confuses people, already tagged.
- Pricing and quoting history. What things cost over time, how quotes vary by size or region, how often they’re accepted. In trades and B2B services this is often the most sought-after number in the sector.
- Hiring and applicant flow. Applications per opening, time to fill, which requirements shrink a candidate pool. Every employer wants this and almost none publish it.
- Inventory, fulfilment and timing data. How long things take, how often they’re late, how seasonality moves. Operational timings are boring internally and interesting externally — that gap is where citations live.
- Your own support response data. Volumes, resolution times, seasonal peaks, channel mix.
- Survey answers you already collect. Signup questions, onboarding forms, exit surveys. You may already hold thousands of responses to a question nobody has ever reported on.
- Public data your industry has but nobody assembled. Records, filings, registers, schedules — scattered across sources and never compiled. Assembling them is origination, and it needs no proprietary access at all.
The pattern to look for: a question people in your field guess at, where a system you own contains the answer.
Three tests a candidate has to pass
Most items on that inventory won’t survive these. That’s fine — you need one or two.
1. Do people argue about it or guess at it? A number is citable in proportion to how unknown it was. If everybody already knows the answer, publishing it changes nothing. If the answer is routinely estimated, disputed, or asked in every forum thread on the subject, you have something. Suppose you run a payroll tool: the share of your customers who file late is a number people speculate about constantly and nobody outside payroll providers can see.
2. Can you publish it without exposing anyone? Covered below, but it’s a gate, not an afterthought. If the number can’t be aggregated safely, its interest is irrelevant.
3. Can you produce it again next year? The test people skip, and the one that separates a nice page from a permanent asset. A repeatable series compounds: each edition earns links, older editions keep theirs, and eventually you’re the standing reference writers check annually. Prefer a duller number you can regenerate to a fascinating one you can’t.
Privacy and aggregation are a prerequisite, not a step
Do this part first, because it can kill the project and you want to know early.
- Aggregate, always. Publish shares, medians, distributions and counts — never rows.
- Threshold small groups. A category with three members isn’t anonymous just because you didn’t name them. Set a minimum group size before you look, and suppress anything under it — including cells that shrink when you slice by two dimensions at once.
- Beware of unique combinations. “Customers in this country, in this industry, of this size” can identify exactly one company even though every field looks harmless alone.
- Check your contracts and privacy policy. What you told customers you’d do with their data governs this, and “we published it as a statistic” is not automatically covered.
- If you can’t anonymise it, drop it. Not “publish it carefully” — drop it and go back to the inventory. There is always another number, and no citation is worth a disclosure.
Doing the analysis honestly
Original data is only an asset if it’s right, and it goes wrong predictably.
Pick the question before you look. Decide what you’re measuring and how, in writing, before you run anything. Otherwise you’ll try fifteen cuts and publish whichever came out interesting — that’s manufacturing a finding, not discovering one.
Report the boring result if that’s the result. “Almost nothing changed year over year” is a real finding and a perfectly citable one. Rerunning an analysis until it yields a headline is very easy to do accidentally.
Show the denominators. A share means nothing without the population it came from. State how many records, over what period, from which subset — and what the subset excludes. Your customers are not your industry, and saying so makes the number stronger, because it tells a citing writer exactly what they’re allowed to claim.
Never round up into a headline you can’t support. If the figure sits near a memorable threshold, resist. The threshold version is what everyone will quote, and it will be wrong.
Name the caveats yourself. Anything you don’t disclose, a critic gets to discover.
Presentation
State the number, the population it describes, and the period — in that order, before any commentary. A writer needs all three to cite you responsibly, and if they have to hunt for the population and the dates, some will guess, and the misquote becomes partly your fault.
Then a short method note, then the breakdowns, then whatever it means. Keep interpretation separate from measurement so somebody can cite the number without buying your opinion about it. Give it one stable page of its own, titled with the metric, so it can be pointed at precisely.
Repeat it
Put next year’s run in the calendar the day you publish. The second edition is far cheaper than the first — the query exists, the method is written, the structure is set — and it’s where the compounding starts. Keep each year at its own address, link them to each other, and the series becomes the default citation because it’s the only continuous record that exists.
What to do next
Pick one number from the inventory that passes all three tests, and write the method down before you query anything. The method note is what makes the number safe for other people to cite — that’s the subject of publishing your methodology.
If nothing on your inventory survives, reread primary sources get cited, summaries get skimmed: you may be able to own something upstream that isn’t a dataset at all. And if your original knowledge is procedural rather than numerical — what happens when you actually do the thing — worked examples earn citations is the format for it.