loader image
Skip to main content

GreenRiver Technology World

AI vs Human Analysts in HR tech Selection

Table of Contents

Every six months we try out a past client HR tech selection request to see how well AI could handle the selection and test the results against our own HRSF ratings.

We then compare with past results to see what improvement, if any, has come about.

The Original Assignment

A UK company with 550 employees (factory and office-based) needed the following:

  • A cloud-based HR system with self-service, performance management, organisation charts and reporting.
  • A payroll module with online payslips and weekly & monthly pay frequencies.
  • A time management module that captures time for factory and office-based employees as well as hybrid and remote workers.
  • A recruitment module that handles aspects such as populating job boards and social media, interview scheduling and routing applications to hiring managers.
  • A learning module (a desirable option at a later date) that would enable them to create their own training and onboarding materials.

We had been asked for a shortlist of five solutions, so we checked in again with the currently popular LLM AI applications. This time, Llama was added to the mix.

For September 2026, we introduced something new:

Which 3 of the named contenders had most AI capability and which 3 had the most compliant published AI governance.

Evaluation and Scoring Method

Using the HRSF ratings, we identified a long list of the best rated 12 products.

  1. For Overall rating, the scoring system was as follows:

    5 points for any product named by the AI appearing in our top 5 choices.
    2 points for any product appearing in our selections rated #6–#12.
    Maximum possible score: 25 points.

  2. For both AI specific Overall scores, the scoring system was as follows:

    In each category we scored 5 points if a selection named in top 3 and 2 points if in our top 10.

The Selections, in Order of Placing

1st — Gemini

Overall AI Features AI Governance
Access People XD Dayforce HiBob
Zellis One Access People XD Dayforce
HiBob HiBob Zellis One
Personio
Dayforce

2nd — ChatGPT

Overall AI Features AI Governance
MHR PeopleFirst MHR PeopleFirst MHR PeopleFirst
Zellis One Employment Hero Employment Hero
Access People XD Access People XD Access People XD
Employment Hero
Element Suite

3rd — DeepSeek

Overall AI Features AI Governance
HiBob ADP Workforce Now ADP Workforce Now
ADP Workforce Now HiBob HiBob
Access People XD Access People XD Ciphr
Workable + Rippling
Ciphr

4th Equal — Manus

Overall AI Features AI Governance
Access People XD Access People XD Sage People
Zellis One Zellis One Access People XD
MHR iTrent Sage People Zellis One
Ciphr
Sage People

4th Equal — Claude

Overall AI Features AI Governance
MHR iTrent HiBob HiBob
Access People XD Ciphr No selection
Ciphr UKG Ready / Pro No selection
UKG Ready / Pro
HiBob

6th — Llama

Overall AI Features AI Governance
Dayforce Workday Workday
Access People XD SAP Success Factors SAP Success Factors
Iris Cascade Hri Dayforce / Oracle HCM Dayforce / Oracle HCM
Ciphr
Sage People

7th — Co-Pilot

Overall AI Features AI Governance
Personio Personio Workday
Employment Hero Rippling SAP Success Factors
HiBob HiBob Oracle Fusion
Rippling
Workday

8th — Perplexity

Overall AI Features AI Governance
Rippling SAP Success Factors Workday
Workday Workday SAP Success Factors
SAP Success Factors Rippling Rippling
Ciphr
HiBob

Overall Score (out of 25)

Previous two results in brackets.

AI Score Previous results
Gemini 17 (12, 8)
ChatGPT 14 (19, 19)
DeepSeek 12 (6, 4)
Llama 10 (–, –)
Manus 10 (19, –)
Claude 9 (19, 8)
Co-Pilot 4 (5, 2)
Perplexity 2 (11, 8)

AI Scores

AI – Features

AI Score
Gemini 9
ChatGPT 7
Claude 7
DeepSeek 6
Manus 4
Llama 2.5
Co-Pilot 2
Perplexity 0

AI – Governance

AI Score
Gemini 12
ChatGPT 7
DeepSeek 7
Manus 9
Claude 5
Llama 0
Co-Pilot 0
Perplexity 0

Commentary

A remarkable result by Gemini and a dire result by Perplexity. ChatGPT showed consistency although with a lower score, and Manus and Claude both lost ground. Newcomer Llama showed promise first time out.

It’s sobering to consider that the winning score was 68% (76% last time), so there’s still a wide margin of error when one relies on AI.

Some selections were out of scope: Oracle, Workday and SAP are not viable considerations for this size of client.

We expected to see products such as Access People XD, Zellis One, HiBob, Dayforce, ADP Workforce Now and UKG Ready / Pro, and these figured well.

There were several good contenders named by the LLMs that just fell outside of our long list for the requirements stated, and Sage People, Ciphr, Personio and MHR iTrent deserve honourable mentions.

Some that certainly would have qualified but didn’t get a mention from any of the AI providers, notably XCD HR, Cezanne HR, Darwinbox and Frontier Software.

I am speculating whether this is because of influence on the AI being caused by presentation of websites or products, rather as SEO caused aberrations in times gone by. It’s hard to see how or why these should have been overlooked.

Both Co-Pilot and Llama had brain fade on the AI selections, naming software that was not in their five contenders. The others seemed to grasp the requirement, except that Claude only committed itself to one selection: HiBob.

One that must be considered, that wouldn’t fit easily with the parameters of the original client request, would be Shapes. Being AI-native, it operates as a “headless” HR tech product, meaning that rather than being divided into modules set by the vendor, the data can be arranged to the client’s own requirements.

Client convenience will almost certainly cause this trend to multiply (subject to legislation on AI-based products not being too restricting), and next time around we expect to see this and similar models appearing.

Conclusions

This round underlines two things. First, AI’s ability to replicate a specialist HR tech shortlist remains inconsistent. Gemini’s strong showing aside, a winning score of 68% (down from 76% last time) means no single AI model is yet reliable enough to replace human analysis outright, though the gap is narrowing for some tools while widening for others.

Second, and perhaps more significant for how we scope future assignments, the exercise surfaced a structural question that sits outside the scoring model altogether: whether the modular, vendor-defined architecture we used to frame this brief will still be the right frame in twelve months’ time.

We mentioned Shapes, and this is a pointer for the future. Rather than a fixed set of modules (HR, payroll, recruitment, learning) bolted together by a vendor, a “headless approach” exposes the underlying data and lets the client or their AI agent assemble the views and workflows they need.

If that model gains traction, our RFPs, Procurement and Demo processes stop being a checklist of modules and instead become a list of data, access controls, and permissions, along with provable governance. This also changes the type of data to be captured, who (or what) can see or act on it, and under what constraints.

Clearly that will have a direct bearing on how we run this comparison next time. The client request becomes a list of capabilities, not modules.

A like-for-like shortlist of “systems” becomes harder to define once the underlying data layer is decoupled from the interface sitting on top of it, and our scoring method may need a parallel track for headless or composable platforms. We’re taking note now, rather than waiting for it to force a redesign of the methodology mid-cycle.

Skip to content
Verified by MonsterInsights