The AI product landscape is moving faster than most buyers can track. New tools appear weekly, vendors make sweeping capability claims, and the marketing noise makes it genuinely hard to separate signal from hype. At huby, our answer to this problem is simple: rigorous, independent evaluation with every decision made in the open.

We publish our full evaluation methodology at huby.ai/methodology. This post walks you through how it works — and why we built it the way we did.

It Starts with Classification

Before we assess anything, we classify the product. We assign it to a product type (for example, conversational AI assistant, code generation tool, or AI-powered search engine) and then identify a more specific product subtype by mapping it against the competitive landscape.

This step matters because the right evaluation criteria for a coding assistant look very different from the right criteria for an enterprise AI search tool. Classification determines which subcategories and assessment factors are activated for that evaluation — and we document and disclose this in the opening section of every report, so you always know the basis of comparison before reading a single score.

A Three-Tier Framework Built for Precision

Our evaluation framework is organized across three tiers:

  • Categories — six broad dimensions that apply across all product types
  • Subcategories — measurable dimensions specific to the product subtype (for example, Access Control Mechanisms within Security)
  • Atomic factors — the specific, verifiable data points being assessed (for example, availability of native multi-factor authentication)
Structural variability by design. The framework is not artificially symmetrical. A category with broader regulatory surface area — like Privacy — carries more subcategories than one with a narrower scope. A complex domain like infrastructure security carries more atomic factors than a more bounded capability. The structure follows the actual shape of the evaluation problem, not a tidy template.

The Six Evaluation Categories

Every product is assessed across six categories, regardless of product type.

Category 1 Quality

How accurately, consistently, and reliably does the product deliver on its stated purpose? We look at output accuracy, user experience design, documentation and developer support, independent benchmarks, and defect rates and resolution times.

Category 2 Use Cases & Pricing

How broad and deep is the supported use case coverage? Does the product offer genuinely differentiated capabilities? We assess ecosystem integration, tier structure, licensing model transparency, and value relative to competitive alternatives — including adoption signals like user growth and retention where data is available.

Category 3 Security

What controls protect the product, its infrastructure, and user data? We consider third-party security audit reports, published CVEs and remediation records, penetration testing outcomes, API security posture, data encryption in transit and at rest, and access control mechanisms.

Category 4 Privacy

How transparent, consistent, and user-centric is the product’s data handling? We assess privacy policy clarity, compliance certifications (SOC 2, GDPR, CCPA/CPRA, HIPAA where applicable), third-party data sharing practices, data retention and deletion controls, and user rights around access, portability, and consent.

Category 5 Sustainability & Reliability

Evaluated across two dimensions: company sustainability — funding status and runway, revenue trajectory, team depth, and market position — and service reliability — uptime commitments, published SLAs, historical incident records, and performance benchmarks. Environmental and social practices are noted where publicly documented.

Category 6 Impact, Ethics & Safety

Does the product and its developer act responsibly toward users and society? We examine the existence and substance of a published ethics policy, bias mitigation practices, model transparency, controls to prevent misuse and harm, and the company’s track record of responding to ethical failures or safety incidents.

Scoring: Evidence-Anchored, 1.0 to 5.0

All scores are reported on a 1.0–5.0 scale at atomic factor, subcategory, and category levels. Each band has a specific evidence requirement — scores are not assigned by feel.

5.0 ExemplaryStrong, independently verified evidence of best-in-class practice. Formal audits, third-party certifications, or peer-reviewed assessments confirm performance.
4.0–4.9 StrongMultiple credible, corroborated sources substantiate the claim. Minor gaps exist but don’t materially undermine the finding.
3.0–3.9 AdequateEvidence is present but limited, mixed, or partially corroborated. Basic standard met, with notable gaps or inconsistencies.
2.0–2.9 WeakMeaningful shortfalls. Claims made by the product owner are not substantiated by independent sources, or documented failures exist.
1.0–1.9 DeficientMaterial failures or absence of the capability or control. Significant risk to users or enterprise buyers.
N/A is not zero. Where evidence is insufficient to score a factor, it is recorded as Not Assessable (N/A) and disclosed in the report. N/A factors are excluded from score calculations entirely rather than treated as zero — preventing evidence gaps from artificially deflating a product’s scores.

Weighted Aggregation — Locked Before Data Collection

Scores aggregate upward through the three-tier hierarchy using a weighted model established during framework design — not a simple average, and not a model that can be adjusted once we’ve seen the results.

Within each subcategory, atomic factors carry defined weights reflecting relative criticality. For example, the presence of multi-factor authentication in an Access Control subcategory is weighted more heavily than audit log granularity, because its absence represents more material risk. Weights within each subcategory sum to 100%. The same logic applies from subcategory to category.

No post-hoc adjustments. Weights are published alongside each evaluation report and cannot be changed after data collection begins. Any change would require a full re-evaluation and a new report version. This means scores cannot be quietly adjusted to favour a result after the fact.

A Five-Tier Source Hierarchy

Not all evidence is equal. huby uses a tiered source framework where higher-tier sources take precedence when sources conflict. All sources cited in a report must be independently accessible via a verifiable URL.

  • Tier 1

    Independent third-party audits and certifications
    SOC 2 reports, CVE databases, named security firm assessments, peer-reviewed research
  • Tier 2

    Established independent journalism and analysis
    Major technology publications, analyst firms, academic institutions
  • Tier 3

    Product owner documentation
    Official privacy policies, terms of service, technical documentation, changelogs
  • Tier 4

    Structured user and practitioner evidence
    Enterprise case studies from named customers, verified professional community posts
  • Tier 5

    Anecdotal community evidence
    Reddit threads, social media posts, informal reviews
huby does not accept payment from product owners in exchange for favorable scores. For newer products without a long public evidence trail, product owners may submit documentation directly — those submissions are identified as such in citations and receive no preferential weighting. Commercial relationships, where they exist, are disclosed in the relevant report.

Evidence Gaps Are Disclosed, Not Buried

Where a product capability cannot be assessed due to insufficient public evidence, we record it explicitly as an evidence gap — we don’t infer from absence, and we don’t assign a low score based on speculation. This is disclosed in the report and results in an N/A score for that factor. Where sufficient public data exists, scores are also contextualised against the product’s competitive set.

Two Report Types for Different Audiences

Public Transparency Report

Covers all six evaluation categories with subcategory-level scores and written assessments. Evidence is cited with verifiable URLs. Available to all huby users at no cost.

Detailed Owner Report

Produced for the product owner. Includes atomic factor-level scores, specific improvement recommendations, and competitive gap analysis. Confidential to the product owner.

Limitations Worth Knowing

huby reports are grounded in publicly available evidence and product owner submissions at the time of evaluation. They do not constitute a formal security audit, legal compliance certification, or financial advisory opinion. Scores reflect the state of the product at the report production date and may not reflect subsequent changes.

Importantly, the absence of evidence for a capability or control does not confirm its absence — it reflects the limits of publicly available information at evaluation time. Where legal, security, or financial interpretation is required, we say so and encourage readers to seek qualified professional advice.

Why This Approach?

AI product evaluation is genuinely difficult. Evidence is scattered, vendor claims are optimistic, and the stakes for enterprise buyers making adoption decisions are real — security exposure, regulatory risk, wasted investment.

We built our methodology to be transparent precisely because we believe the credibility of our findings depends on readers being able to assess the validity — and the limitations — of our work for themselves. The framework is public. The weights are published before data collection begins. The sources are cited and verifiable. The gaps are disclosed rather than papered over.

That is what independent evaluation actually means.

Ready to see the methodology applied to a real product? Browse our AI product evaluations on huby.

Explore AI Product Reports →