Research Toolkit / Delivering Insights
Delivering insights is the process of minimizing communication errors and creating a narrative arc that leaves your audience inspired to build upon your work.
The understanding of others sets the ceiling on the perceived brilliance of your ideas.
Three models, read as a sequence: reproduce what prior work found, show where your evidence departs from it, then reconcile the two. The reconciliation is the contribution.
| Predictor | M1 · Replicate the prior result | M2 · Diverge your result | M3 · Reconcile both, together |
|---|---|---|---|
| X — the predictor prior work credited | 0.45* | 0.12 | 0.28 |
| Z — the predictor you introduce | — | 0.38* | 0.22 |
| X × Z — the two interacting, which is the mechanism | — | — | 1.13* |
Read the top row across: what prior work credited to X survives on its own (0.45*), collapses once Z is in the model (0.12), and turns out to have been the two acting together all along (1.13*). Three questions to ask of any such table: does M2 challenge the inference of M1, does M3 reconcile the tension rather than merely add a term, and is the reconciliation both theoretically motivated and empirically explained? · Illustrative values.
“A presentation table is a figure made of numbers. Highlight the cells that represent the shift in belief.”
Each beat of the arc has a picture. Before fitting a single model, you should be able to sketch the three figures that carry the story — and if a figure needs a thousand words of caption to land, it is not yet doing its job.
The test of whether you understand your own argument is whether you can draw it. If you cannot sketch the picture for a beat, you do not yet have that beat.
One job only: establish that there is something to explain. No controls. A reader cannot wave it away, so it motivates the question.
Show the setting where the competing accounts would leave different fingerprints — and that they do.
The gap between the comfortable estimate and the design-based one. That distance is the finding.
Ask a room how they would handle selection and most will say they would add controls. Controls reach only the part of the confound you observed. The design — a lottery, an instrument, a discontinuity, a within-unit falsification — is what reaches the part you did not.
The wrong way runs a battery of alternative specifications and reports that the headline coefficient survives all of them. But a results section in which every check holds is not reassuring — it is implausible, and a careful reader knows it. Real data are noisier than that.
Robustness checks should not be exercises in results preservation but, rather, sincere probing of the conditions that warrant belief updating.
A wall of confirmations reads as a wall built to keep the reader out.
The better way pre-commits to specifications under which the effect should weaken or vanish if the mechanism is the right one, and treats misbehaviour as information. Add tests of the mechanism’s side-predictions so the analysis can find the boundary of the effect, not only its centre. Tie any reassurance to the specific doubt the design actually leaves open, never to a generic battery.
You cannot see what you wrote. The sentence that is obvious to you, because you know what you meant, is the one a referee will stumble over — and you are structurally unable to notice it. Run a structured review against your own draft before anyone else does. It surfaces the standard objections in their standard form: you claim to do this but the table does not show it; the paragraph comparing the ideal design to the actual one is missing; this variable’s construction is not stated.
It is often wrong about what to do. It is reliably useful about what you are avoiding. Let it find the objection; you decide the answer.
The pivot from “yeah, yeah, yeah” to “but” is the load-bearing joint. These are the four ways it fails, each with the symptom that gives it away.
Overturning a consensus the paper never stated fairly. No prior is loaded, so nothing moves — and the experts you skipped past are now your hostile referees. Symptom: an introduction that argues before it describes.
The phenomenon is established, accounts are motivated, and then it stops — confirmation offered as contribution. Symptom: a final table whose preferred column says nothing the first column did not.
Real models ordered by convenience or software default rather than as beats. Model 3 does not answer what Model 1 raised; controls accrete without an argument. Symptom: a reader who follows every row and still cannot say what the point is.
The introduction promises a pivot the evidence cannot deliver; the design never closes off the comfortable account, so the “but” is asserted rather than earned. Symptom: prose saying “we address this concern” where a column should be doing the addressing. The most dangerous of the four — it survives until the referee checks.
Justify your empirical context as a strategic asset. Use this checklist to diagnose the strength of your identification strategy. After Al-Ubaydli, List & Suskind (2017), The science of using science: Towards an understanding of the threats to scaling experiments, NBER Working Paper 23032.
How were actors sorted into this setting?
Who is missing? Is survival bias clouding the inference?
Does the setting mirror real decision-making incentives?
Does the mechanism operate predictably as the system grows?
Describe the experiment you would run with a magic wand. Describe what you actually did. Explain why the gap is defensible and how you address it. Reviewers respect scholars who explicitly compare their actual design to the ideal; identifying threats to inference yourself is a signal of methodological expertise.
“You probably won’t have — and don’t need — a bulletproof or gold-standard design. But you must demonstrate that you understand the shortcomings and have a plan for addressing them.”
Do not rush the a-ha or impress with verbal velocity.
If the back row cannot read it, it does not exist.
Slides are evidence, not a script.
Do not answer an empirical question with theory if you have a figure; and do not answer a theoretical question with data.
“Never show a regression coefficient until you have shown the raw data that justifies it. Seeing a plot prepares the audience to see the coefficient.”