The Calibration Call

A short research conversation with someone who has done the role or hired for it, to learn what great and failed performance look like before you define it.

What is The Calibration Call?

You’re the CEO, a career engineer, and you’ve just sold the first half a million in revenue, which means now you’re hiring salespeople for the first time. You know your way around sales, but you have no idea how to hire salespeople, let alone what’s reasonable to expect from someone unable to draw upon the magical powers of being a founder. You’re lost, unsure of how to define the role, create a rubric, or evaluate candidates.

You rewatch that clip from Glengarry Glen Ross about coffee being for closers, ready to prove you can close salespeople, not just deals. [1] You stick a note to the bottom of your monitor, two words in Sharpie: “Must close.” But the best salesperson you know is the guy at Best Buy who sold you that extended care plan. “Maybe we don’t need salespeople,” you think. “If I can figure it out, other engineers can figure it out.” After spending a bunch of money, losing a good engineer or two, and getting no closer to hitting that sales goal, you’re back at hiring salespeople. To evaluate candidates, you realize you need something to compare them to. You need a standard.

To develop that standard, you will make a series of Calibration Calls, short research conversations with experts who fit in three groups:

  1. People who are actively succeeding in the job you’re hiring for
  2. People who have successfully hired for the role you’re hiring for
  3. People who have worked with people great at the role you’re hiring for

You may also need to do a series of routing calls with your network to get access to people in those three buckets.

The point is not to outsource your judgment, but to calibrate what great looks like from people who know what the fuck they’re talking about.

The difference between a great calibration call and a poor one depends on three things: who you talk to, what you ask them, and what you do with their answers.

Getting to the right people

Selection

But don’t only call the most famous person you know. In any field where luck plays a real part in who succeeds — and startups require plenty of luck — the celebrated winners may not be the most reliable sources of signal on what works. [2] Two or three currently practicing experts are likely to beat one legend. [3]

Use peer nomination for routing, but it’s a proxy for expertise. Most functions (role + seniority) you hire for in startups don’t have transparent scoreboards. There is no international database ranking VPs of Marketing. Instead you rely on peer nomination, a fancy label for asking people who they think is good. To increase your chances of talking to a real expert, look for people who:

  • Are doing the job right now, not a decade ago.
  • Are at a scale and stage relevant to yours.
  • Can point to tangible outcomes.

Ask routing questions. You won’t always know five top performers in the role you’re hiring for, but someone you already know might. Ask, “Who do you know who is best in class at this?” Then ask, “Who do you know who would know a lot of people who are best in class at this?” This is called snowball sampling. [4] My favorite, a question particularly good for experts: “who do you go to when you’re stuck?” And ask routing questions to weak ties (acquaintances, neighbors, people you used to work with). Weak ties reach clusters your close contacts can’t.

Talk with two to three people actively doing the job, and up to two from each of the other two buckets. That’s four to seven calls, plus however many routing calls it takes to reach them. As long as you find high quality people in each bucket, the returns on additional calls after 2-3 are marginal, and the research supports that: accuracy after two opinions is ~27% and rises to 33% after eight. [5] Quick refresh on the buckets:

  1. People who are actively succeeding in the job you’re hiring for
  2. People who have successfully hired for the role you’re hiring for
  3. People who have worked with people great at the role you’re hiring for

Talk with people from different social clusters. You graduated from MIT in 2010, so you ask five friends from your class. Five perspectives, right? Wrong. You got one. [6] This is called homophily or “sampling from a homogeneous seed.” [6] Correlated advisors from the same network yield accuracy gains of 7% vs. independent ones at 27%. [7] Five people from one network is one opinion. Five from five networks is five.

Access

People say yes more than you think. [8] You must stop believing excellent people are too busy to speak with you. People underestimate the chance of a yes by as much as half. [8] When we ask for help, we weigh the cost of someone helping us, but miss that refusing to help is a cost most people aren’t willing to pay.

Embrace routing calls if you need to. In the best case scenario, everyone you need to talk to in buckets one through three are close first-degree connections. Reality is rarely a best case scenario. Thankfully, routing calls fix that. Use the questions in routing questions above.

Do the work to make it easy to say yes. You receive two emails. One is full of typos and flattery, ending with a vague ask. The second is well-researched and well-written, closing with a specific ask. You’d archive the first and take the second. Do your research. Write the email yourself. Ask for 20-45 minutes: less and you won’t be taken seriously, more and they’ll feel burdened. [9]

Consider The Ben Franklin Effect. [10] In his autobiography, Franklin recounts gaining the favor of a peer legislator. He’d heard the man had a rare book, and he asked to borrow it. Lending the book was a reasonable request, one simple enough for the man to fulfill. [10] The man had never spoken to him in the House. After the book came back, he did. They were friends until he died. [11] The paragraph ends on the principle: “He that hath once done you a kindness will be more ready to do you another than he whom you yourself have obliged.”

Get people on the phone or face-to-face. People say yes 34 times more often to a face-to-face request than to the same request by email. [12] Use an email or a text to get someone on the phone, but don’t cold-email or text people the questions you need answered.

Running the call

It is hard for an expert to explain something to a novice. You’d think that if you know something, you’d know how to explain it. But the literature disagrees. [13] [14] [15] What follows isn’t superfluous. It is the price of learning what you need to hire a stunning colleague.

Ask for permission to record the call or take great notes. Permission is legally required in many states and it’s good ethics. A transcript lets you compile specific behaviors, outcomes, competencies, expectations, and questions to ask candidates. Without a transcript, take great notes, because you will not remember it. People taking notes while listening capture about a third of the important information and one tenth of the most specific stuff. [16] The specifics are the entire reason you’re doing the call.

What to ask

Be deliberate in the questions you ask. Plan them. James Spradley spent his career learning how outsiders could interview insiders about worlds they don’t belong to. In his 1979 The Ethnographic Interview, he suggests a wide variety of questions and techniques for eliciting information from interviewees. Here are a few of my favorites:

  1. Grand tour questions: Get people to describe the whole territory, encouraging them to “ramble on and on.” There are four subtypes: typical (“Describe a typical day”), specific (“Tell me what you did yesterday, from the time you got to work until you left”), guided (“Could you show me around?”), and task-related (ask them to walk you through how they do a task).
  2. Mini-tour questions: If a “grand tour” question is asking someone what a typical day looks like, and they answer in part calling on prospects, a mini tour then asks what calling on prospects looks like. The outcome should be a list of activities.
  3. Example questions: Take one thing the person said and ask for a specific example.
  4. Personal vs. cultural phrasing: Instead of asking “Can you describe a great closing call?” ask “What do salespeople do in a closing call?” People, strangely, are much worse at describing themselves than the role. [17]

Lead with questions about failure. People aren’t good at telling you what their standard for excellence is, even if they’re experts. [18] [19] But if you ask an expert all the ways people screw a thing up, or what people who’ve failed at the job all have in common, you’re much more likely to get useful data.

Discount predictions. High performers will confidently look at someone’s LinkedIn and tell you whether they think your candidate will be successful. They’ll do it with fluency, confidence, and specificity. They believe competence at doing grants competence at judging. It does not. Daniel Kahneman and Gary Klein, two psychologists who disagreed on many things, concluded together that experts are often called upon to “make judgments in areas in which they have no skill.” [20] They called it fractionated expertise and believed it was the rule, not the exception. [20] James Shanteau sorted expert fields by whether the experts in them are any good. “Personnel selectors”, hiring managers, made the bad list.

Seek observations. Ask what successful people do. “If a camera crew followed you around for a week, what would we see? What does the work involve?” Find out how the best people succeed. For example, ask “How would a great salesperson handle a stalled enterprise deal?”

How to ask

How you ask matters as much as what you ask, and nearly all of it comes down to making it easier for people to answer. The first thing I learned as an executive coach was not to start questions with “why” unless I meant to challenge someone. [9] Why reads as criticism and puts people on the defensive; what and how don’t. [9] “How does that work?” “What makes that counterproductive?” Ask follow-ups, so they know you’re listening. [21] Give people an escape hatch, telling them they can change their answers later, and they’ll say more. [21] Finally, repeat their language back in their words, not yours. [9]

Keep asking until you understand. You’re going to spend a lot of time on calibration calls not understanding things, and you’re going to want to be helpful. “Hey, so this is what I think we need for this job. Sound right?” It feels like preparation. It feels like giving them context. In almost any other professional conversation, you’d be right. But here it just hands them your answer to react to, narrowing the field of information you’re likely to collect. [9]

What to weigh

Distinguish between craft fluency and performance. A lot of people can “talk the talk” without being able to “walk the walk.” Interviews with candidates are much better at assessing someone’s craft fluency than at assessing performance. [22] As you calibrate, ask experts how to distinguish between the two.

Put sample sizes in context with each person you call. Having assessed more than 5,000 investment managers, Graham Duncan suggests you assume “80% of people are not calibrated on the Michael Jordan” of whatever you’re hiring for. [23] [24] He relies on this question:

What’s your sample size of people in the role in which you knew Jane?… If we call [the best you’ve seen] a ‘100’, where’s Jane right now on a 1-100? [24]

When speaking with people who hired for the role rather than doing it themselves, beware that they will tell you what they believe good looks like, which is distinct from what good is. [25] To separate belief from reality, ask “What do you screen for now that you didn’t screen for in the past?”

Afterwards

Once you’ve finished all the calls related to a role, work through each of them all at once to integrate what you learned into your hiring process. Different people’s accounts are likely to contradict each other, and you’ll have to reconcile what you heard. Incorporate specific outcomes and skills you heard into your MOC, and try any interview questions and tactics you gathered in your interview process. What you heard about the first few months should inform the expectations you set on day one. And write a short thank you note to everyone for their time.

To hire someone, you need to evaluate candidates. To evaluate candidates, you need to compare them against a standard. If you’ve never done the job before, you don’t have a standard. Calibration calls are how you develop a standard. Calibration calls are relevant to two Management Craft competencies:

  1. Defining roles: how can you define a role if you’ve never done it or spoken to people who’ve done it?
  2. Knowing what great looks like: you need a standard of greatness to compare candidates to.

A calibration call is the management equivalent of user research. A good leader will obsessively talk to customers to figure out what to build. The same leader will hire a VP of Sales having never talked to a great one and wonder why it didn’t work out.

Andy Grove reduced all management to four activities: (1) collect information, (2) convey information, (3) make decisions, and (4) be a role model. [26] Calibration calls are, at first, about collecting information so you can decide who to hire. And later, they’re about conveying information in the documents you write next. If the assumptions you build those documents on are vague, you can’t expect your people to deliver.

Use the Calibration Call when you’re hiring for a function, stage, or level you don’t already understand from direct experience. It’s especially useful before your first executive hire, a new function, or a role where you or your team can describe the pain but not a standard of excellence.

Hiring is not a linear process. It loops back on itself. You draft, you realize you missed something, you call someone, you revise, rinse, and repeat. So there isn’t a single “right” moment to start making calibration calls. I recommend them when you’re writing an MOC, when you can’t make an outcome measurable or can’t name what separates good from great. It’s a great tool when you’re designing an interview loop or when you don’t know what to test for in a role. Any of these are fine. What matters most is that you use calibration calls to create and refine the standard you’ll use to evaluate candidates.

You can skip calibration calls when you already have a standard. If you’ve done the job yourself, or hired for it several times and tracked how those hires turned out, you have usable reference points and don’t need to use someone else’s. Just make sure not to trick yourself. A founder who sold the first half million knows how to sell. They do not know what a great VP of Sales looks like.

Research base

There is no academic research proving that calibration calls work, only self-report and anecdotes. Everything cited on this page is evidence about tactics or parts, not the whole: how people answer requests, how badly experts model beginners, how much a second opinion is worth.

Operator base

I recommend it because I have seen the first-hand benefit of conducting calibration calls for more than a decade. It is one of the most frequently prescribed practices I recommend to clients I coach.

And I’m not alone in recommending it. Liz Wessel (First Round Capital), [27] Keith Rabois (Khosla Ventures), [28] Brian Armstrong (Coinbase), [29] Elad Gil (angel investor), [28] Sam Altman (OpenAI now; YC when he endorsed), [30] and Brian Chesky (Airbnb) [28] have all commented on the usefulness of various forms of the calibration call.

The practice is old. It’s also convergent: people keep arriving at it independently, without a name and without talking to each other. What they don’t agree on is how many calls to conduct. [27] [28] [29] [30]

Only recently, however, did I discover a name for the practice. In late 2025 and again in early 2026, Liz Wessel wrote about the concept, calling them “calibration calls.” [27] She didn’t claim to invent it, but since I’ve never seen anyone else describe this specific technique as a “calibration call”, I’m attributing it to her unless someone shows me otherwise.

The term “calibration call” has been circulating for a long time, but it’s always been used to refer to a different kind of call than the one written about here. Recruiters often recommend a “calibration call” with a hiring manager to make sure they’re on the same page. HR will conduct “calibration calls” with managers during performance review season, [31] and different interviewers on an interview panel often do the same. [32] Someone has information, the other person needs it.

Finally, a bit of trivia. The root word of calibrate is caliber. The Oxford English Dictionary declares it a borrowing from the French calibre, the bore of a gun barrel. [33] Where that French word came from, OED says, is uncertain: the Arabic qālib, a mould for casting metal, or a cognate of qalaba, to turn, has been suggested. [33] And in a footnote, Carl August Friedrich Mahn, who wrote the 1864 Webster’s, conjectured the Latin quâ librâ: of what weight? [33] A bore, a mould, a turn on a lathe, a known weight: every root of the word is a thing someone made so the next thing could be checked against it. A caliber is a standard. To calibrate is to go and find one.

1
James Foley, dir., Glengarry Glen Ross, screenplay by David Mamet, New Line Cinema, 1992. Blake’s speech, the one that ends “Coffee’s for closers only,” was written for the film and is not in Mamet’s stage play.
https://en.wikipedia.org/wiki/Glengarry_Glen_Ross_(film)
2
Jerker Denrell and Chengwei Liu, “Top Performers Are Not the Most Impressive When Extreme Performance Indicates Unreliability,” Proceedings of the National Academy of Sciences 109, no. 24 (2012): 9331-9336.
https://doi.org/10.1073/pnas.1116048109
3
Richard P. Larrick and Jack B. Soll, “Intuitions About Combining Opinions: Misappreciation of the Averaging Principle,” INSEAD Working Paper 2003/09/TM, 2003.
https://flora.insead.edu/fichiersti_wp/inseadwp2003/2003-09.pdf
4
Jyoti Shankar Tripathy, Akhilesh Singh, and Deepanjali Tripathy, “The Double-Edged Sword of Consecutive and Snowball Sampling: Practical Utility Versus Methodological Compromise,” Indian Journal of Psychological Medicine 48, no. 1 (2026): 81-84.
https://doi.org/10.1177/02537176251405469
5
Ilan Yaniv and Maxim Milyavsky, “Using Advice from Multiple Sources to Revise and Improve Judgments,” Organizational Behavior and Human Decision Processes 103, no. 1 (2007): 104-120.
https://doi.org/10.1016/j.obhdp.2006.05.006
6
Miller McPherson, Lynn Smith-Lovin, and James M. Cook, “Birds of a Feather: Homophily in Social Networks,” Annual Review of Sociology 27 (2001): 415-444.
https://doi.org/10.1146/annurev.soc.27.1.415
7
Ilan Yaniv, Shoham Choshen-Hillel, and Maxim Milyavsky, “Spurious Consensus and Opinion Revision: Why Might People Be More Confident in Their Less Accurate Judgments?,” Journal of Experimental Psychology: Learning, Memory, and Cognition 35, no. 2 (2009): 558-563. The figures are gains in accuracy over the pre-advice estimate, not accuracy rates.
https://doi.org/10.1037/a0014589
8
Francis J. Flynn and Vanessa K. B. Lake, “If You Need Help, Just Ask: Underestimating Compliance with Direct Requests for Help,” Journal of Personality and Social Psychology 95, no. 1 (2008): 128-143.
https://doi.org/10.1037/0022-3514.95.1.128
9
Duke Initiative on Survey Methodology, “Tipsheet: Interviewing Elites,” accessed August 3, 2026.
https://dism.duke.edu/files/2020/05/Tipsheet-Elite_Interviews.pdf
10
Meg Jay, The Defining Decade, updated ed., Twelve, 2021, ch. “Weak Ties.”
https://www.hachettebookgroup.com/titles/meg-jay-phd/the-defining-decade/9781538754238/
11
Benjamin Franklin, The Autobiography of Benjamin Franklin, ed. John Bigelow, Lippincott, 1900, pp. 216-217.
https://www.gutenberg.org/files/20203/20203-h/20203-h.htm
12
M. Mahdi Roghanizad and Vanessa K. Bohns, “Ask in Person: You’re Less Persuasive Than You Think over Email,” Journal of Experimental Social Psychology 69 (2017): 223-226.
https://doi.org/10.1016/j.jesp.2016.10.002
13
Pamela J. Hinds, “The Curse of Expertise: The Effects of Expertise and Debiasing Methods on Prediction of Novice Performance,” Journal of Experimental Psychology: Applied 5, no. 2 (1999): 205-221.
https://doi.org/10.1037/1076-898X.5.2.205
14
Colin Camerer, George Loewenstein, and Martin Weber, “The Curse of Knowledge in Economic Settings: An Experimental Analysis,” Journal of Political Economy 97, no. 5 (1989): 1232-1254.
https://doi.org/10.1086/261651
15
Mitchell J. Nathan, Kenneth R. Koedinger, and Martha W. Alibali, “Expert Blind Spot: When Content Knowledge Eclipses Pedagogical Content Knowledge,” in Proceedings of the Twenty-Third Annual Conference of the Cognitive Science Society, 2001.
16
Kenneth A. Kiewra, Tiphaine Colliot, and Junrong Lu, “Note This: How to Improve Student Note Taking,” IDEA Paper #73, IDEA Center, September 2018.
https://www.ideaedu.org/idea_papers/note-this-how-to-improve-student-note-taking/
17
James P. Spradley, The Ethnographic Interview, Holt, Rinehart and Winston, 1979, pp. 86-88.
https://archive.org/details/ethnographicinte0000spra
18
Gary Klein, Sources of Power: How People Make Decisions, MIT Press, 1998, ch. 4.
https://mitpress.mit.edu/9780262534291/sources-of-power/
19
Anders Ericsson and Robert Pool, Peak: Secrets from the New Science of Expertise, Houghton Mifflin Harcourt, 2016, ch. “The Gold Standard.”
https://www.harpercollins.com/products/peak-anders-ericssonrobert-pool
20
Daniel Kahneman and Gary Klein, “Conditions for Intuitive Expertise: A Failure to Disagree,” American Psychologist 64, no. 6 (2009): 515-526.
https://doi.org/10.1037/a0016755
21
Alison Wood Brooks and Leslie K. John, “The Surprising Power of Questions,” Harvard Business Review, May-June 2018.
https://hbr.org/2018/05/the-surprising-power-of-questions
22
Harry Collins and Robert Evans, Rethinking Expertise, University of Chicago Press, 2007.
https://press.uchicago.edu/ucp/books/book/chicago/R/bo5485769.html
23
Graham Duncan, “The Playing Field,” grahamduncan.blog, 2018.
https://grahamduncan.blog/the-playing-field/
24
Graham Duncan, “What’s Going On Here, With This Human?,” grahamduncan.blog, 2021.
https://grahamduncan.blog/whats-going-on-here/
25
David C. McClelland, “Testing for Competence Rather Than for ‘Intelligence,’” American Psychologist 28, no. 1 (1973): 1-14.
https://doi.org/10.1037/h0034092
26
Andrew S. Grove, High Output Management, Random House, 1983, p. 98.
https://www.amazon.com/High-Output-Management-Andrew-Grove/dp/0679762884
27
Liz Wessel, “Hiring for a New Role/Function for Your First Time? Do This First,” X, February 18, 2026; earlier, in shorter form, as an X post, August 8, 2025.
https://x.com/lizwessel/status/2024206318324920443
28
Elad Gil, High Growth Handbook, Stripe Press, 2018, ch. 4, “Building the Executive Team,” pp. 122, 125. The passage at p. 125 is Gil’s interview with Keith Rabois, who credits Brian Chesky with the technique.
https://growth.eladgil.com/book/chapter-4-building-the-executive-team/
29
Brian Armstrong, “How to Hire Executives,” Medium, August 8, 2017.
https://barmstrong.medium.com/how-to-hire-executives-e2ee8e05cad3
30
Sam Altman, “How to Hire,” blog.samaltman.com, September 23, 2013.
https://blog.samaltman.com/how-to-hire
31
Korn Ferry, “What HR Leaders Need to Know About Performance Calibration,” kornferry.com, accessed August 3, 2026.
https://www.kornferry.com/insights/featured-topics/employee-experience/hr-leaders-and-performance-calibration
32
Pin, “Interview Debrief and Calibration: Align Your Hiring Panel,” pin.com, 2026.
https://www.pin.com/blog/interview-debrief-guide/
33
“calibre | caliber, n.,” Oxford English Dictionary, Oxford University Press, etymology, accessed August 3, 2026.
https://www.oed.com/dictionary/calibre_n?tab=etymology

Keep learning

Defining rolesCompetencyHow to go from a vague hiring need to a clearly defined and documented standard your team can use to consistently evaluate candidates.Knowing what great looks likeCompetencyWhen hiring for roles you've never done, you need to learn what the best people in that role do to make them so great before you start interviewing.Turning whispers into standardsCompetencyThe things you most want (and don't want) in teammates are the hardest to name. Turn gut feelings into criteria you can actually interview for.The Competency StackToolA five-layer standard for what a hire must clear: the role, the level, the company's values, its stage, and the team they're joining.The MOCToolThe MOC is the document you write before you talk to a single candidate: the Mission the role exists to serve, the Outcomes you will hold the hire to, and the Competencies it takes to deliver them. Get it right and the same page becomes your interview guide, your onboarding plan, your ninety-day scorecard, and your grounds for letting someone go if it comes to that.The Terrain TestToolA hiring screen for stage fit: has the candidate recently succeeded under the same ambiguity, structure, and support this role will demand?