The AI product landscape is moving faster than most buyers can track. New tools appear weekly, vendors make sweeping capability claims, and the marketing noise makes it genuinely hard to separate signal from hype. At huby, our answer to this problem is simple: rigorous, independent evaluation with every decision made in the open.
We publish our full evaluation methodology at huby.ai/methodology. This post walks you through how it works — and why we built it the way we did.
It Starts with Classification
Before we assess anything, we classify the product. We assign it to a product type (for example, conversational AI assistant, code generation tool, or AI-powered search engine) and then identify a more specific product subtype by mapping it against the competitive landscape.
This step matters because the right evaluation criteria for a coding assistant look very different from the right criteria for an enterprise AI search tool. Classification determines which subcategories and assessment factors are activated for that evaluation — and we document and disclose this in the opening section of every report, so you always know the basis of comparison before reading a single score.
A Three-Tier Framework Built for Precision
Our evaluation framework is organized across three tiers:
- Categories — six broad dimensions that apply across all product types
- Subcategories — measurable dimensions specific to the product subtype (for example, Access Control Mechanisms within Security)
- Atomic factors — the specific, verifiable data points being assessed (for example, availability of native multi-factor authentication)
The Six Evaluation Categories
Every product is assessed across six categories, regardless of product type.
Category 1 Quality
How accurately, consistently, and reliably does the product deliver on its stated purpose? We look at output accuracy, user experience design, documentation and developer support, independent benchmarks, and defect rates and resolution times.
Category 2 Use Cases & Pricing
How broad and deep is the supported use case coverage? Does the product offer genuinely differentiated capabilities? We assess ecosystem integration, tier structure, licensing model transparency, and value relative to competitive alternatives — including adoption signals like user growth and retention where data is available.
Category 3 Security
What controls protect the product, its infrastructure, and user data? We consider third-party security audit reports, published CVEs and remediation records, penetration testing outcomes, API security posture, data encryption in transit and at rest, and access control mechanisms.
Category 4 Privacy
How transparent, consistent, and user-centric is the product’s data handling? We assess privacy policy clarity, compliance certifications (SOC 2, GDPR, CCPA/CPRA, HIPAA where applicable), third-party data sharing practices, data retention and deletion controls, and user rights around access, portability, and consent.
Category 5 Sustainability & Reliability
Evaluated across two dimensions: company sustainability — funding status and runway, revenue trajectory, team depth, and market position — and service reliability — uptime commitments, published SLAs, historical incident records, and performance benchmarks. Environmental and social practices are noted where publicly documented.
Category 6 Impact, Ethics & Safety
Does the product and its developer act responsibly toward users and society? We examine the existence and substance of a published ethics policy, bias mitigation practices, model transparency, controls to prevent misuse and harm, and the company’s track record of responding to ethical failures or safety incidents.
Scoring: Evidence-Anchored, 1.0 to 5.0
All scores are reported on a 1.0–5.0 scale at atomic factor, subcategory, and category levels. Each band has a specific evidence requirement — scores are not assigned by feel.
| 5.0 | ExemplaryStrong, independently verified evidence of best-in-class practice. Formal audits, third-party certifications, or peer-reviewed assessments confirm performance. |
| 4.0–4.9 | StrongMultiple credible, corroborated sources substantiate the claim. Minor gaps exist but don’t materially undermine the finding. |
| 3.0–3.9 | AdequateEvidence is present but limited, mixed, or partially corroborated. Basic standard met, with notable gaps or inconsistencies. |
| 2.0–2.9 | WeakMeaningful shortfalls. Claims made by the product owner are not substantiated by independent sources, or documented failures exist. |
| 1.0–1.9 | DeficientMaterial failures or absence of the capability or control. Significant risk to users or enterprise buyers. |
Weighted Aggregation — Locked Before Data Collection
Scores aggregate upward through the three-tier hierarchy using a weighted model established during framework design — not a simple average, and not a model that can be adjusted once we’ve seen the results.
Within each subcategory, atomic factors carry defined weights reflecting relative criticality. For example, the presence of multi-factor authentication in an Access Control subcategory is weighted more heavily than audit log granularity, because its absence represents more material risk. Weights within each subcategory sum to 100%. The same logic applies from subcategory to category.
A Five-Tier Source Hierarchy
Not all evidence is equal. huby uses a tiered source framework where higher-tier sources take precedence when sources conflict. All sources cited in a report must be independently accessible via a verifiable URL.
-
Tier 1
Independent third-party audits and certifications
SOC 2 reports, CVE databases, named security firm assessments, peer-reviewed research -
Tier 2
Established independent journalism and analysis
Major technology publications, analyst firms, academic institutions -
Tier 3
Product owner documentation
Official privacy policies, terms of service, technical documentation, changelogs -
Tier 4
Structured user and practitioner evidence
Enterprise case studies from named customers, verified professional community posts -
Tier 5
Anecdotal community evidence
Reddit threads, social media posts, informal reviews
Evidence Gaps Are Disclosed, Not Buried
Where a product capability cannot be assessed due to insufficient public evidence, we record it explicitly as an evidence gap — we don’t infer from absence, and we don’t assign a low score based on speculation. This is disclosed in the report and results in an N/A score for that factor. Where sufficient public data exists, scores are also contextualised against the product’s competitive set.
Two Report Types for Different Audiences
Public Transparency Report
Covers all six evaluation categories with subcategory-level scores and written assessments. Evidence is cited with verifiable URLs. Available to all huby users at no cost.
Detailed Owner Report
Produced for the product owner. Includes atomic factor-level scores, specific improvement recommendations, and competitive gap analysis. Confidential to the product owner.
Limitations Worth Knowing
huby reports are grounded in publicly available evidence and product owner submissions at the time of evaluation. They do not constitute a formal security audit, legal compliance certification, or financial advisory opinion. Scores reflect the state of the product at the report production date and may not reflect subsequent changes.
Importantly, the absence of evidence for a capability or control does not confirm its absence — it reflects the limits of publicly available information at evaluation time. Where legal, security, or financial interpretation is required, we say so and encourage readers to seek qualified professional advice.
Why This Approach?
AI product evaluation is genuinely difficult. Evidence is scattered, vendor claims are optimistic, and the stakes for enterprise buyers making adoption decisions are real — security exposure, regulatory risk, wasted investment.
We built our methodology to be transparent precisely because we believe the credibility of our findings depends on readers being able to assess the validity — and the limitations — of our work for themselves. The framework is public. The weights are published before data collection begins. The sources are cited and verifiable. The gaps are disclosed rather than papered over.
That is what independent evaluation actually means.
Ready to see the methodology applied to a real product? Browse our AI product evaluations on huby.