When people ask “What’s the best AI chatbot?”, they’re usually expecting a single answer.

After evaluating six of the most widely used AI chatbots across six dimensions of trust, we can give you one — but we also want to show you why stopping there would miss the point.

Google Gemini leads our first AI chatbot evaluation with a Trust Score of 4.3 out of 5.0, followed closely by Claude (4.2) and ChatGPT (4.1). But the more important story is what lies beneath those numbers.

That’s why huby doesn’t believe AI products should be judged by one number alone.

Why We Started This Evaluation

AI chatbots are no longer experimental tools. People use them to write documents, generate code, analyze data, learn new skills, and make decisions that matter. Yet most publicly available comparisons focus on a narrow set of capability benchmarks or a handful of headline features.

We wanted to answer a broader question: how should users evaluate whether an AI product deserves their trust?

Looking Beyond Model Performance

Model capability is only one part of the picture. For this evaluation, we assessed six products — ChatGPT, Claude, DeepAI, DeepSeek, Google Gemini, and Grok — across six dimensions:

  • Quality — How well does the product perform for its intended purpose, consistently and reliably?
  • Privacy — How does it handle user data, and how clearly does it communicate its practices?
  • Security — What protections are in place for users and their information?
  • Sustainability & Reliability — How dependable is the service over time, and how sound is the company behind it?
  • Use Cases & Pricing — Does it deliver practical value across different types of users and contexts?
  • Impact, Ethics & Safety — How responsibly does the product address misuse, transparency, and broader societal considerations?

The Trust Score is a weighted composite across these six categories. Weights are set before data collection begins and published alongside our full evaluation methodology — they cannot be adjusted after the fact to favour any result.

The Results

#1
Google Gemini
4.3
Trust Score · Strong
Claude
4.2
Trust Score · Strong
ChatGPT
4.1
Trust Score · Strong
DeepSeek
3.5
Trust Score · Adequate
DeepAI
3.4
Trust Score · Adequate
Grok
2.9
Trust Score · Weak

The Trust Scores above show where each product lands overall. Now here’s what the detail reveals.

huby AI Chatbot Leaderboard — scores for ChatGPT, Claude, DeepAI, DeepSeek, Google Gemini, and Grok across Quality, Privacy, Security, Sustainability, Use Cases, and Impact, Ethics & Safety
Category-level scores across all six trust dimensions — on huby’s 1.0–5.0 scale

The top three sit within 0.2 of each other in Trust Score, but their profiles look very different. Claude leads on Privacy (4.5) — the highest single-category score in the entire evaluation. Google Gemini leads on Use Cases & Pricing (4.5). ChatGPT is the most consistent, with no category below 3.7. If Privacy is your priority, Claude is the clear choice. If breadth of use cases matters most, Gemini edges ahead.

The gap between the top three and the rest is significant. Nearly 1.5 points separates Gemini (4.3) from Grok (2.9) in Trust Score — not a minor difference. Grok’s Security (2.6) and Impact, Ethics & Safety (2.6) scores represent meaningful concerns for enterprise or sensitive-use deployments.

Impact, Ethics & Safety is where every product has the most room to improve. It produces the lowest category scores across the evaluation — a signal the industry as a whole has work to do.

What Surprised Us

The biggest lesson from this evaluation wasn’t about which product ranked highest. It was how differently products perform across categories — and how misleading a Trust Score alone can be.

Claude and Gemini are separated by just 0.1 in Trust Score (4.2 vs 4.3), but their category profiles tell a different story: Claude is the clear privacy leader while Gemini leads on practical utility. For most individual users, that distinction matters more than the composite gap.

Evaluating AI is inherently multidimensional. Reducing everything to a single Trust Score risks hiding differences that matter. A user making a decision about enterprise deployment needs to know where the strengths and gaps actually are — not just a composite number.

Trust Is More Than Performance

Research from Pew and others consistently shows that AI adoption continues to grow while public confidence remains mixed. That disconnect suggests trust depends on more than impressive capability demonstrations.

People increasingly care about questions that benchmark scores don’t answer:

  • How is my data used — and by whom?
  • Can I rely on the output for consequential decisions?
  • How transparent is the company about what the product can and can’t do?
  • What safeguards exist against misuse?
  • How has the company responded when things have gone wrong?

These questions deserve objective, evidence-based answers. That’s what huby is built to provide.

Why Transparency Matters

At huby, we believe an evaluation is only as valuable as the methodology behind it. That’s why we publish not only our Trust Scores but the complete evaluation framework used to produce them — including the criteria, the evidence sources, and how scores are weighted and aggregated.

Transparency is not a feature of our evaluation process — it is one of its core principles. We expect the methodology to evolve as AI evolves, and we welcome feedback from users, researchers, and product teams. Read the full methodology at huby.ai/methodology.

Where We Go From Here

This chatbot evaluation is the beginning of something larger, and there are two things worth flagging as you use these results.

Trust Scores will change. AI is one of the fastest-moving fields in technology. The products we evaluated today will ship new versions, update their privacy practices, earn new security certifications, or change their pricing models. When they do, our evaluations will be updated to reflect it. huby Trust Scores represent the state of a product at the time of evaluation — not a permanent verdict.

Coding agents are next. Our next evaluation category is AI coding agents — tools like GitHub Copilot, Cursor, and others that are increasingly central to how software gets built. We’ll apply the same six-dimension framework, with subcategories and weighting adapted to what matters most for that product type. If you work in software development and want to weigh in on what we should prioritize, we’d welcome the input.

More broadly, our goal is to build a consistent, transparent Trust Score across all major categories of AI products — so that users, organizations, and developers can make better-informed decisions as the landscape continues to shift.

Because the future of AI won’t be determined solely by what AI can do. It will also be shaped by how much people trust the products they choose to use.

Explore the full chatbot leaderboard and dive into the Trust Scores behind each category.

See the Full Results on huby.ai →

Evaluation methodology: huby.ai/methodology