Every six months we try out a past client HR tech selection request to see how well AI could handle the selection and test the results against our own HRSF ratings.
We then compare with past results to see what improvement, if any, has come about.
The Original Assignment
A UK company with 550 employees (factory and office-based) needed the following:
- A cloud-based HR system with self-service, performance management, organisation charts and reporting.
- A payroll module with online payslips and weekly & monthly pay frequencies.
- A time management module that captures time for factory and office-based employees as well as hybrid and remote workers.
- A recruitment module that handles aspects such as populating job boards and social media, interview scheduling and routing applications to hiring managers.
- A learning module (a desirable option at a later date) that would enable them to create their own training and onboarding materials.
We had been asked for a shortlist of five solutions, so we checked in again with the currently popular LLM AI applications. This time, Llama was added to the mix.
For September 2026, we introduced something new:
Which 3 of the named contenders had most AI capability and which 3 had the most compliant published AI governance.
Evaluation and Scoring Method
Using the HRSF ratings, we identified a long list of the best rated 12 products.
- For Overall rating, the scoring system was as follows:
5 points for any product named by the AI appearing in our top 5 choices.
2 points for any product appearing in our selections rated #6–#12.
Maximum possible score: 25 points. - For both AI specific Overall scores, the scoring system was as follows:
In each category we scored 5 points if a selection named in top 3 and 2 points if in our top 10.
The Selections, in Order of Placing
1st — Gemini
| Overall | AI Features | AI Governance |
|---|---|---|
| Access People XD | Dayforce | HiBob |
| Zellis One | Access People XD | Dayforce |
| HiBob | HiBob | Zellis One |
| Personio | ||
| Dayforce |
2nd — ChatGPT
| Overall | AI Features | AI Governance |
|---|---|---|
| MHR PeopleFirst | MHR PeopleFirst | MHR PeopleFirst |
| Zellis One | Employment Hero | Employment Hero |
| Access People XD | Access People XD | Access People XD |
| Employment Hero | ||
| Element Suite |
3rd — DeepSeek
| Overall | AI Features | AI Governance |
|---|---|---|
| HiBob | ADP Workforce Now | ADP Workforce Now |
| ADP Workforce Now | HiBob | HiBob |
| Access People XD | Access People XD | Ciphr |
| Workable + Rippling | ||
| Ciphr |
4th Equal — Manus
| Overall | AI Features | AI Governance |
|---|---|---|
| Access People XD | Access People XD | Sage People |
| Zellis One | Zellis One | Access People XD |
| MHR iTrent | Sage People | Zellis One |
| Ciphr | ||
| Sage People |
4th Equal — Claude
| Overall | AI Features | AI Governance |
|---|---|---|
| MHR iTrent | HiBob | HiBob |
| Access People XD | Ciphr | No selection |
| Ciphr | UKG Ready / Pro | No selection |
| UKG Ready / Pro | ||
| HiBob |
6th — Llama
| Overall | AI Features | AI Governance |
|---|---|---|
| Dayforce | Workday | Workday |
| Access People XD | SAP Success Factors | SAP Success Factors |
| Iris Cascade Hri | Dayforce / Oracle HCM | Dayforce / Oracle HCM |
| Ciphr | ||
| Sage People |
7th — Co-Pilot
| Overall | AI Features | AI Governance |
|---|---|---|
| Personio | Personio | Workday |
| Employment Hero | Rippling | SAP Success Factors |
| HiBob | HiBob | Oracle Fusion |
| Rippling | ||
| Workday |
8th — Perplexity
| Overall | AI Features | AI Governance |
|---|---|---|
| Rippling | SAP Success Factors | Workday |
| Workday | Workday | SAP Success Factors |
| SAP Success Factors | Rippling | Rippling |
| Ciphr | ||
| HiBob |
Overall Score (out of 25)
Previous two results in brackets.
| AI | Score | Previous results |
|---|---|---|
| Gemini | 17 | (12, 8) |
| ChatGPT | 14 | (19, 19) |
| DeepSeek | 12 | (6, 4) |
| Llama | 10 | (–, –) |
| Manus | 10 | (19, –) |
| Claude | 9 | (19, 8) |
| Co-Pilot | 4 | (5, 2) |
| Perplexity | 2 | (11, 8) |
AI Scores
AI – Features
| AI | Score |
|---|---|
| Gemini | 9 |
| ChatGPT | 7 |
| Claude | 7 |
| DeepSeek | 6 |
| Manus | 4 |
| Llama | 2.5 |
| Co-Pilot | 2 |
| Perplexity | 0 |
AI – Governance
| AI | Score |
|---|---|
| Gemini | 12 |
| ChatGPT | 7 |
| DeepSeek | 7 |
| Manus | 9 |
| Claude | 5 |
| Llama | 0 |
| Co-Pilot | 0 |
| Perplexity | 0 |
Commentary
A remarkable result by Gemini and a dire result by Perplexity. ChatGPT showed consistency although with a lower score, and Manus and Claude both lost ground. Newcomer Llama showed promise first time out.
It’s sobering to consider that the winning score was 68% (76% last time), so there’s still a wide margin of error when one relies on AI.
Some selections were out of scope: Oracle, Workday and SAP are not viable considerations for this size of client.
We expected to see products such as Access People XD, Zellis One, HiBob, Dayforce, ADP Workforce Now and UKG Ready / Pro, and these figured well.
There were several good contenders named by the LLMs that just fell outside of our long list for the requirements stated, and Sage People, Ciphr, Personio and MHR iTrent deserve honourable mentions.
Some that certainly would have qualified but didn’t get a mention from any of the AI providers, notably XCD HR, Cezanne HR, Darwinbox and Frontier Software.
I am speculating whether this is because of influence on the AI being caused by presentation of websites or products, rather as SEO caused aberrations in times gone by. It’s hard to see how or why these should have been overlooked.
Both Co-Pilot and Llama had brain fade on the AI selections, naming software that was not in their five contenders. The others seemed to grasp the requirement, except that Claude only committed itself to one selection: HiBob.
One that must be considered, that wouldn’t fit easily with the parameters of the original client request, would be Shapes. Being AI-native, it operates as a “headless” HR tech product, meaning that rather than being divided into modules set by the vendor, the data can be arranged to the client’s own requirements.
Client convenience will almost certainly cause this trend to multiply (subject to legislation on AI-based products not being too restricting), and next time around we expect to see this and similar models appearing.
Conclusions
This round underlines two things. First, AI’s ability to replicate a specialist HR tech shortlist remains inconsistent. Gemini’s strong showing aside, a winning score of 68% (down from 76% last time) means no single AI model is yet reliable enough to replace human analysis outright, though the gap is narrowing for some tools while widening for others.
Second, and perhaps more significant for how we scope future assignments, the exercise surfaced a structural question that sits outside the scoring model altogether: whether the modular, vendor-defined architecture we used to frame this brief will still be the right frame in twelve months’ time.
We mentioned Shapes, and this is a pointer for the future. Rather than a fixed set of modules (HR, payroll, recruitment, learning) bolted together by a vendor, a “headless approach” exposes the underlying data and lets the client or their AI agent assemble the views and workflows they need.
If that model gains traction, our RFPs, Procurement and Demo processes stop being a checklist of modules and instead become a list of data, access controls, and permissions, along with provable governance. This also changes the type of data to be captured, who (or what) can see or act on it, and under what constraints.
Clearly that will have a direct bearing on how we run this comparison next time. The client request becomes a list of capabilities, not modules.
A like-for-like shortlist of “systems” becomes harder to define once the underlying data layer is decoupled from the interface sitting on top of it, and our scoring method may need a parallel track for headless or composable platforms. We’re taking note now, rather than waiting for it to force a redesign of the methodology mid-cycle.
