Research

Committee vs. Algorithm: What We Learned from 200 Retroactive Audits

By Crewpath Team  · 

Committee versus algorithm decision comparison

When we ran a retroactive scoring pass against 200 completed engagements from a mid-size management consulting firm — scoring each engagement as if the model had been running at the time of the original committee decision — the divergences were more instructive than the agreements. The model and the committee agreed on the top match in about 58% of cases. In the remaining 42%, the question was not simply who was right. The question was what each approach was optimizing for, and whether the firm knew.

What the retroactive audit methodology looks like

A retroactive audit applies the scoring model to historical data: engagement records, consultant rosters as they existed at the time, post-project client ratings, and the committee's actual deployment decisions. The model produces a ranked list for each historical engagement. The auditor then compares the model's top recommendation to the committee's choice, and for cases where they diverge, examines the dimension-level scores to understand why.

This is not a judgment on whether the committee made a "mistake" in any given case. Many committee decisions that diverge from the model's recommendation are defensible on factors the model didn't have access to: informal partner knowledge, context about a specific client relationship, a consultant's stated career development goals, circumstances that weren't documented in any system. The retroactive audit is a diagnostic, not a verdict.

Where committee judgment consistently added value

In engagements involving a new client relationship where the originating partner had specific knowledge about the client's culture and expectations, committee overrides of the model's top recommendation were frequently correct. The model's client chemistry dimension draws from historical engagement data — which, for a new client, is zero. A partner who had done pre-sales work with a prospective client over six months had substantially better information about fit than any dimension score derived from roster data alone.

Complex engagements requiring a consultant who was technically strong but atypically matched — a consultant moving into a new service line by design, or a pairing that made sense for development reasons the firm had articulated internally — also showed a pattern of committee decisions being sensible in ways the model couldn't see. When firms made deliberate development allocations, the scoring model flagged those consultants as lower-ranked by domain expertise. The committee was right to override.

Where the algorithm consistently outperformed

The divergence category that produced the clearest pattern of committee underperformance was familiarity bias. In 34 of the 200 engagements, the committee selected a consultant who ranked 4th or lower in the model's output. Post-project feedback scores for those engagements ran meaningfully below average. Domain expertise scores for the selected consultant were substantially lower than for the model's top-ranked choice. In most of these cases, the selected consultant had worked with the originating partner before — the relationship, not the fit, drove the decision.

Utilization-driven allocation errors were the second consistent pattern. In 22 cases, the committee selected a consultant who ranked 1st on utilization pressure (most in need of billable work) but 5th or lower on domain expertise. These were cases where bench pressure appears to have overridden fit evaluation. Post-project outcomes for this cluster were the worst in the audit — not dramatically, but measurably below both the average and the cases where the committee selected the model's top recommendation.

The cases where both approaches were incomplete

In roughly 15% of the divergent cases, neither the model's recommendation nor the committee's choice produced strong post-project outcomes. These engagements had structural problems that no allocation decision could solve: understaffed for scope, unclear deliverable definition, client-side misalignment that emerged after the project started. Allocation quality is one driver of engagement success among several. We don't suggest that a perfect scoring model would prevent all staffing-related project problems, because many project problems aren't staffing problems.

What firms actually changed after seeing the audit results

The most common change after retroactive audit is not a wholesale shift to model-first allocation. It's a recalibration of the committee's role: narrowing from "who should we consider for this engagement" to "does the model's top recommendation make sense, and if we're overriding it, what is the specific reason." In that structure, the committee's time is spent on the cases that genuinely require judgment, rather than on discovery work the model handles faster and more consistently.

Firms also typically make changes to their client chemistry data collection after seeing audit results. The dimension that showed the widest variance in model reliability was chemistry — because the underlying feedback data was sparse or inconsistent for many client relationships. Knowing this, firms become more deliberate about capturing post-engagement feedback as a systematic process rather than an optional close-out step.

The design intent behind the audit feature

Crewpath's retroactive audit mode exists precisely because firms should not have to take a model's efficacy on faith. Running the model against 12-24 months of historical engagements and comparing scored output to actual decisions gives leadership a concrete basis for calibrating trust in the model, adjusting weights to match firm-specific values, and identifying the specific decision categories where human judgment adds clear value versus categories where the model's consistency is the right answer.