The take-home assignment is now the least informative part of most interview loops. It was never a great instrument, but it worked well enough when producing a competent deck, a clean model or a working prototype took a weekend of real effort. The artifact was evidence of the effort. Now the artifact can be produced in an afternoon by someone with mediocre instincts and a good tool, and it will look roughly as polished as the one produced by the person you actually want. Every candidate clears the bar, so the bar tells you nothing.
This is the part of the AI conversation that gets least attention and matters most to anyone doing the hiring. The interesting shift is not that machines write code or first drafts. It is that the cost of producing plausible work has collapsed while the cost of being wrong has not moved at all. Hiring for judgment used to be a luxury you exercised at the senior level after you had confirmed someone could do the job. It is now the whole job of the interview.
Judgment is two skills, and most loops test neither
When people say they want to hire for judgment they usually mean something vague and admirable. It is more useful split in two.
The first half is selection: deciding what is worth producing at all. When output was expensive, scope control happened naturally, because nobody had the hours to build the thing nobody asked for. Cheap production removes that friction. I now see teams generating more analysis, more variants, more documentation and more dashboards than anyone reads, and mistaking the volume for progress. The person who is worth hiring is the one who looks at a request and says that three of these five deliverables do not change any decision, so let us not make them.
The second half is detection: noticing when the output is wrong. Confidently wrong work is harder to catch than obviously bad work, and the current generation of tools produces confidently wrong work with excellent formatting. A junior analyst who does not understand a market will write something hesitant and vague, and you will spot it. A model will write something fluent and specific and slightly false. Detecting that requires a person who knows what the right answer should roughly look like before they read the output.
Most interview processes test a third thing entirely — the ability to produce. That skill has not become worthless, but it has become common, and it is no longer the thing that separates candidates. I wrote a while back about where AI actually earns its keep in an operating company, and the same logic applies to the people you put around it. The tool amplifies whatever judgment is already in the room. Hire poorly and you have simply bought a faster way to be wrong.
Interviewing for selection: give them too much and watch what they cut
The exercise I have come to trust is deliberately overscoped. Hand the candidate a realistic situation with more possible work in it than anyone could do in the time allowed, and make clear — explicitly, so they are not guessing at the rules — that you are not expecting all of it. Then ask what they would do first, what they would not do at all, and why.
What you learn is immediate. Weak candidates try to cover everything, because covering everything feels safe and they have spent their careers being rewarded for thoroughness. Strong candidates cut aggressively and defend the cuts. The defence is where the signal lives. "I would skip the competitive teardown because whichever answer it produces, we still take the same next step" is a sentence that tells you more than any portfolio.
The follow-up matters as much as the exercise. Ask what would have to be true for their priority to be wrong. People with real judgment answer that quickly and without defensiveness, because they have already asked themselves the question. People without it hear a challenge and start reselling the original answer. I use a similar move in diligence, where the aim is to find the question whose answer predicts trouble rather than the question that produces the most information. Interviews reward the same economy.
Interviewing for detection: hand them something plausible and wrong
The second exercise is simpler and, in my experience, the single most useful forty minutes in any loop. Give the candidate a finished-looking artifact in their domain — a memo, a model, a plan, a piece of analysis — that contains a small number of seeded errors. Make the errors the kind that survive a skim: a reasonable-sounding assumption that contradicts another one three pages later, a figure carried forward incorrectly, a conclusion that does not follow from the evidence above it, a source that is cited for something it does not say.
Do not tell them the errors are there. Ask them to review it as though a team member had sent it for sign-off before it went to a client or a board. Then watch.
Some candidates copyedit. Some restructure. Some praise it. The ones you want read for internal consistency, and they tend to find the contradiction before they find the typo. Just as telling is what they do when they are unsure — whether they flag the uncertainty plainly or quietly let it pass to avoid looking slow. That instinct, the willingness to say "I do not think this number is right and I cannot yet tell you why," is close to the core of what an AI-assisted team needs and is very hard to teach after the fact.
One warning about this exercise. It only works if the reviewer is fluent enough in the domain to have priors. That is why I am sceptical of the emerging idea that you can hire people whose skill is directing tools rather than doing the work. Verification is a craft skill. You cannot catch a wrong margin assumption if you have never built a margin assumption. The right hire is not someone who has stopped doing the work; it is someone who has done enough of it to know when the output smells wrong.
The résumé signal that still holds
Credentials have always been a weak proxy and are weaker now, because the polish they used to signal is available to everyone. The signal I still weight heavily is accountability — whether the candidate has ever personally carried the consequences of a call they made.
People who have owned outcomes talk about their work differently. They volunteer what went wrong without being asked. They can name a belief they held two years ago and no longer hold, and explain what changed their mind. People who have only ever produced deliverables talk about scope, effort and stakeholders, and describe outcomes as things that happened near them. Neither type is a bad person. Only one of them has been trained by consequences, which is the only reliable trainer of judgment I have encountered.
So I ask for the decision, not the project. What did you decide, who disagreed, and how did you find out you were right or wrong? If the candidate has never been in a position to find out, that is useful information too — it tells you what you are actually buying.
The junior problem, which is real
There is an obvious objection to all of this. If judgment comes from accountability and accountability comes from experience, and the entry-level tasks that used to build experience are now automated, where does the next generation of judgment come from?
I do not think the answer is to stop hiring early-career people. I think it is to hire them for a different quality and then give them the missing exposure deliberately. Through a board I serve on at the FIU Honors College I spend time around students who are near the start of this, and the ones who stand out are not the most polished — they are the ones who revise fastest when they are shown they are wrong. That is the trait to interview for at the junior level. Show them their own mistake mid-interview and watch what happens in the next ten minutes. I have written before that students are visibly bad at things while executives have learned to hide it, and that visibility is an asset in a hiring process, not a liability.
Then put the junior hire in front of consequences early and in small doses. Let them own a real recommendation that someone will act on. Judgment is not transferred by observation; it is built by being wrong in public and having to fix it.
Treat the hire as a decision you will review
The uncomfortable part of hiring for judgment is that you are exercising the same faculty you are trying to measure, and you will get it wrong at a rate you would find unacceptable in any other part of the business. The only real defence is to write down, before the offer goes out, what you believed about this person and what you expect to see within their first few months. Then go back and read it. A hiring decision is exactly the kind of call that deserves a proper post-mortem rather than a shrug, and the process improves only if you keep the record honest.
We are heading into a period where nearly everyone's work product looks good. What will distinguish teams is whether someone in the room can tell the difference between good-looking and correct, and whether that person had the standing to say so before it shipped. Hire for that, and protect it once you have it.
