
A few months ago, we explained how we run performance reviews with Claude: connecting our tools via MCP, encoding each evaluation criterion into a skill, and letting the model read months of evidence no human had time to read. That post was about the setup. This one is about what we learned after actually using it.
For context, our performance reviews are built on templates per role: one for Product Managers, one for Tech Leads, one for Software Engineers, and specific ones for marketing and HR. Each template defines the aspects we evaluate, and each item gets a rating from A to D depending on whether the person exceeds, meets, or falls short of expectations.
Being a development company, a big part of what we evaluate is technical. Code quality and the bug rate of what gets shipped. Whether an engineer avoids introducing unnecessary technical debt. For Tech Leads, whether they can make sound architecture decisions and explain technical topics to clients in a way non-technical people understand. Among many other things.
On top of that, we evaluate soft skills too: being communicative, raising doubts, and asking for help when needed. AI usage. The quality of written communication (mistakes, formatting, clarity). Proactivity in suggesting improvements. Participation in internal meetings.
We covered the full criteria in a previous article.
We connected Google Drive, Slack, Linear, and the rest of our stack, and wrote detailed skills for each criterion that can be verified through documentation. Whether a developer has been proactive suggesting improvements shows up in internal meeting notes. Whether an engineer answers the daily Slack message, and with what level of detail, is right there in the channel. We ran the reviews with Fable, and the results on this front were genuinely good.
The model reads everything, across the entire review period, and comes back with concrete examples: this person raised this risk in this meeting, this report flagged this deviation, this issue was well documented. No recency bias, no "I think I remember them doing that once". For pulling evidence out of six months of activity, it beats any human reviewer, simply because no human reads all of it.
The problem appears one step later, when you have to decide what the evidence means.
Take a real example. We expect Tech Leads to provide insights on the projects they run: raise risks, question decisions, flag things before they become problems. The AI checks the meeting notes, finds two occasions where the Tech Lead did exactly that, and takes it as proof that the expectation was met. Rating: meets expectations, evidence attached.
But was it met? Maybe those two occasions were the only ones in six months, on a project full of situations where a Tech Lead should have raised their hand. The AI has no way of knowing how many opportunities existed and were missed. You can ask it to count occurrences, and it will. What it can't do is weigh whether that count is enough for the responsibility this person carries. Two insights from a junior developer might be a great sign. Two insights from a Tech Lead over half a year might be a problem.
The same happens with severity. The AI flags a mistake in a client report, but is it a typo or something that damaged trust with the client? It lists both with the same neutral tone. Deciding how much something matters requires context the tools don't contain: the expectations you set for that person, the difficulty of their project, the conversations that happened outside any system.
We think this is the interesting lesson, because it applies to many other processes where we've been introducing AI. Architecture decisions in development. Evaluating incoming project proposals. Anywhere the hard part is not collecting information but exercising judgment: weighing trade-offs, applying experience, deciding what's acceptable and what isn't.
In all these cases the pattern is the same. The model does the mechanical reading better than we ever could, and then presents everything with the same confident flatness, whether it's a minor detail or a red flag. The calibration is on us.
We'll keep using AI for performance reviews. The time it saves gathering evidence is too big to give up, and reviews backed by actual data are fairer than reviews backed by memory.
But we're adjusting the weights. The AI's output is the starting point, not the conclusion. Peer feedback, client feedback, and the manager's own perception carry more weight in the final rating than they did in our first automated cycle, precisely because those sources come with judgment built in. A colleague who tells you "I expected more from them on this project" is doing something the model can't.

We thought Linear AI was just another tech distraction, until we actually tried it. Find out why this new update is a secret superpower for Product Managers, not just developers.
Leer el artículo
How we integrate agents like Claude and Copilot into our workflow using a rigorous Research, Plan, and Implement framework to ensure speed without sacrificing architectural excellence.
Leer el artículo
What does a great software engineer look like today? A look inside our updated review templates and the new AI criteria we use to evaluate our team.
Leer el artículo