Agent Simulation & Evaluation
Hiring for the Team You Already Have
Think about the last person your team hired who did not work out.
Not the one who couldn't do the job. That case is easy. Everyone saw it, the code didn't compile, the numbers didn't add up, and the parting was clean. I mean the other one. The one who was clearly smart, clearly capable, checked every box on paper, and somehow made the whole team a little worse. Meetings got tense. The quiet engineer got quieter. Two people who used to riff together stopped. Nobody could point to a single thing that was wrong, because nothing was wrong with the person. Something was wrong with the fit.
We almost never see that coming. And I have started to think we no longer have a good excuse.
The two kinds of fit
When you hire, you are really asking two questions.
The first is technical fit. Can this person do the work? We are good at this. We have interviews, take-home problems, system design rounds, references, whole industries built around measuring it. It is not perfect, but it is a real measurement of a real thing.
The second is social fit. Will this person actually mesh with the specific group of people already in the room? Not "are they nice," not "culture fit" as a vague vibe, but the concrete question: how will this one human change the chemistry of these exact ten humans doing their exact daily work?
We pretend to measure this. We do a "culture interview," someone gets a gut feeling, we call it a day. But be honest about what that is. It is a guess dressed up as a process. And the reason it stays a guess is simple and, until recently, unarguable. You cannot run the experiment. You cannot hire someone for three months, watch the team, and then un-hire them and try the next candidate under identical conditions. The one thing you would need to actually know is the one thing you can never do.
So we shrug. We say fit is unmeasurable, an art, a feel. We measure the thing we can measure, technical skill, and then we cross our fingers on the thing that quietly decides whether the team thrives.
I want to argue that the second thing is no longer unmeasurable. You cannot run the experiment on real people. But you can simulate it.
Start with a claim about the models
Here is the foundation, and I want to state it plainly because everything rests on it.
Modern language models are genuinely good at emotional intelligence. Not in a mystical sense. In a practical, testable sense. Give a model a transcript of how a person speaks and writes, and it will read them with unsettling accuracy. How they handle disagreement. Whether they hedge or commit. Whether they take up all the air in a room or wait. Whether criticism lands on them as information or as threat. We read each other this way too, from tone and word choice and what a person does with a pause. The models do it from the same signal, and they do it well.
If that is true, and I believe it is, something follows that is bigger than it first looks. An interview is enough.
An interview is a dense sample of a person. In an hour of real conversation, a person shows you how they think under mild pressure, what they get excited about, how they treat someone with less power than them, where their patience runs out. A model that can read emotional signal can take that hour and build a persona.
Let me be careful with that word. A persona is a working model of a person. Not their soul, not a copy, not a claim to have captured who they really are. It is a structured description, rich enough that you can ask it, "given this situation, how would this person likely respond?" and get an answer that leans the way the real person would lean. The way a good novelist holds a character firmly enough to know they would never say a certain line. That is the bar. Not perfect prediction. Reliable leaning.
Now put the persona to work
Here is the part that turns hiring on its head.
A team has ten people. It wants to hire an eleventh. HR forwards a hundred candidates. Today that hundred gets brutally filtered by resume, and maybe five get real interviews, and of those you extend one offer, and you learn whether the fit was right roughly ninety days too late.
Try it the other way.
You interview all hundred. From each interview you build a persona. Then you build a simulation of your actual team, the real ten, doing their real daily work. The standup where the same person always runs long. The design review where two people always disagree and a third always smooths it over. The Friday crunch when things get short. This is not a fantasy office. It is your office, modeled from how your ten people actually behave, because the same reading that builds a candidate's persona can build theirs.
Then you drop each simulated candidate into that simulated team, one at a time, and you watch.
You watch what the group does. Does the quiet engineer speak less when candidate 34 is in the room? Does the team's disagreement get sharper and meaner, or sharper and more productive? Does candidate 71, brilliant on paper, quietly recentre every conversation on themselves until the others stop offering ideas? You are not scoring the candidate in isolation. You are scoring the group with that candidate in it, against the group without them. You are measuring the delta. You are measuring the thing everyone knows matters and nobody can measure today.
I find this genuinely thrilling, and I want to say why without overselling it. We are not inventing a new virtue to hire for. Social fit was always the quiet decider. We just finally have a way to look at it before we commit, instead of discovering it in the wreckage three months in.
Where the honesty lives
Now the hard part, because this only earns your trust if I say it clearly.
You are predicting, not knowing. A persona is a model of a person, and a model is always less than the thing it models. The real candidate will have a day the interview never saw. They will grow, or curdle, in ways no hour of conversation could forecast. The simulated team is a model too, and it will smooth over the strange, specific, load-bearing details that make real people surprising. Run the same simulation twice and you may get two different Fridays.
So this does not tell you what will happen. It tells you what is likely to happen, which is a weaker claim and a far more useful one than the nothing we have now. Weather forecasts are models of the sky, and no one thinks the forecast is the weather, and yet we would not dream of planning without it. That is the right frame. Not an oracle. A forecast for the one question we used to answer with a coin flip.
The core claim survives all of that caution. This is now possible, and it is worth building. The models can read people. An interview is a rich enough sample. A persona is reliable enough to lean the right way. And a simulation lets you run the experiment you were never allowed to run on the humans themselves.
The person on the other side
There is one more thing, and it is the part I keep coming back to.
The tenth person on that team, the quiet one who gets quieter around the wrong hire, never gets a say today. Nobody asks them, because there is no way to ask. We hire over their heads and hope. What this really does is give the team you already have a voice in who joins it next. Not a veto, not a vibe, a measured signal about what the addition does to the whole.
We spend so much effort measuring whether a person is good enough for the team. This finally lets us ask the quieter, kinder question underneath it.
Is the team going to be good for them, too.