Is Your AI Hiring Tool Science or Snake Oil? Key Takeaways From Hire Ground Live#3
By Tanisha ·
AI hiring tools are getting harder to ignore. They can screen CVs, assess interviews, surface candidates, and even explain why someone appears to be a strong match.
But a good-looking demo does not tell you whether the tool actually works. That was the question at the centre of Hire Ground Live Episode 3, held on September 3: Is Your AI Hiring Tool Science or Snake Oil?
Hosted by Desiree Goldey, the conversation brought together Kate Young, Business Psychologist and Founder of The Hiring Science Studio, and Mayur Macwan, Founder of RoundOne AI.
Kate brought a psychological and assessment-science perspective to the discussion, focusing on how hiring assessments should be designed, tested, and validated. Mayur brought the perspective of an AI hiring founder building and testing these systems in real recruitment environments.
Rather than another conversation about how exciting AI is, the episode focused on a more useful question: What should hiring teams actually ask before trusting an AI hiring tool?
Watch The Full Hire Ground Live Episode
Before getting into the key takeaways, watch the full conversation to hear Kate and Mayur unpack AI assessment science, validity, reliability, bias monitoring, and what hiring teams should demand from AI vendors.
Here’s the full episode of Hire Ground Live Episode 3 on LinkedIn:
Below, we have highlighted some of the key takeaways and excerpts from the conversation.
What Is That AI Score Actually Measuring?
Getting a score from an AI hiring tool can make the process feel objective. But as Kate explained, the score means very little unless you know what went into producing it.
Was the tool measuring communication, planning, technical skills, or simply how much a candidate spoke? What questions were asked, what did the candidate say, and how was that response processed? Those details matter far more than the final number.
As Kate emphasized during the conversation, a measurement is only useful when you understand what it is actually measuring. A shortlist is not automatically a good shortlist just because an algorithm created it.
Building The Tool Is The Easy Part
It has become surprisingly easy to build an AI hiring product. Give an LLM a CV and job description, ask it to compare the two, and you can have a working prototype very quickly.
The harder part is proving that the output means anything. As Mayur discussed, the real test is not whether an AI system can generate a score. It is whether that score holds up when you put it against real candidates and real hiring decisions.
- Does the score hold up with real candidates?
- Does it measure the skill you intended to measure?
- Most importantly, does it tell you anything useful about how someone will perform in the job?
That is where proper testing comes in.
Reliability Does Not Equal Validity
The conversation also unpacked two terms that are often thrown around together: reliability and validity. Reliability is about how stable a measurement is. Validity asks whether you are measuring the right thing in the first place. An AI tool can consistently return the same score and still be consistently wrong.
Take communication as an example. If an assessment rewards candidates simply because they give longer answers, it may appear consistent while actually measuring how much someone talks rather than how well they communicate.
This was one of the important distinctions Kate brought into the conversation: consistency alone does not make an assessment meaningful. A reliable measurement still needs to demonstrate that it is measuring what it claims to measure.
“Bias-Free” Is Not A Useful Promise
The phrase “bias-free AI” also came under fire during the discussion. Kate’s point was simple: AI tools are built by people, so decisions about data, scoring, prompts, and assessment criteria can introduce bias.
The more realistic goal is bias mitigation. Teams need to monitor outcomes, look for differences between groups, and investigate what might be causing them.
As the discussion highlighted, fairness is not something a hiring tool can simply establish once and then forget about. Teams need to keep monitoring the system after deployment and investigate unexpected outcomes.
Your Best Employees Should Not Become The Algorithm
There is another problem with using past hiring decisions as training data. If a company has historically hired people from similar universities, companies, or backgrounds, an AI trained on those patterns may simply reproduce them.
The better question is: what actually makes someone successful in this role?
If strong planning or communication skills predict performance, those are worth assessing. Where someone went to college or which company appears on their CV may not be.
As Mayur’s perspective from building RoundOne showed, real-world testing can help expose whether an AI system is actually identifying useful signals or simply reproducing existing hiring patterns.
So, How Should You Evaluate An AI Hiring Tool?
Kate suggested starting with two straightforward questions:
- “Can I see the technical manual?”
- “Can I speak to your scientist?”
Those questions get to the heart of the issue. Hiring teams should be able to understand how a system works, what evidence supports its assessments, and who is responsible for validating those claims.
Mayur also shared a practical example from a RoundOne pilot. A client took around 50 CVs from a role they had already filled and compared RoundOne’s results with the shortlist created by their recruitment team.
That kind of exercise can tell you much more than a polished demo. You get to see how the tool performs against a real hiring decision rather than an example carefully prepared for a sales call.