The AI product landscape is moving faster than most users can track. New tools appear weekly, some vendors make sweeping capability claims while others go unnoticed. At huby, our answer to this problem is simple: rigorous, independent evaluation with every decision made in the open.
We publish our full evaluation methodology at huby.ai/methodology. This post walks you through how it works — and why we built it the way we did.
It Starts with Classification
Before we assess anything, we classify the product. We assign it to a broad product type (for example, software engineering, creative toolkit, AI Agent, etc) and then identify a more specific product subtype by mapping it against the competitive landscape (e.g. coding assistant, multi-modal, knowledge management).
This step matters because the right evaluation criteria for a coding assistant look very different from the right criteria for an enterprise AI search tool. Classification determines which subcategories and assessment factors are activated for that evaluation — and we document and disclose this in the opening section of every report, so you always know the basis of comparison before reading a single score.
A Three-Tier Framework Built for Precision
Our evaluation framework is organized across three tiers:
- Categories — six broad dimensions that apply across all product types
- Subcategories — measurable dimensions underlying a particular category (for example, Access Control Mechanisms within Security)
- Atomic factors — the specific, verifiable data points being assessed related to a specific subcategory (for example, availability of native multi-factor authentication under Access Control Mechanism)
The Six Evaluation Categories
Every product is assessed across six categories, regardless of product type.
1. Quality
How accurately, consistently, and reliably does the product deliver on its stated purpose? We look at output accuracy and consistency, user experience design, documentation and developer support, independent benchmarks, and defect rates and resolution times.
2. Use Cases & Pricing
How broad and deep is the supported use case coverage? Does the product offer genuinely differentiated capabilities? We assess ecosystem integration, tier structure, free access provisions, licensing model transparency, and value relative to competitive alternatives — including adoption signals like user growth and retention where data is available.
3. Security
What controls protect the product, its infrastructure, and user data? We consider third-party security audit reports, published CVEs and remediation records, penetration testing outcomes, static and dynamic code analysis results, API security posture, data encryption in transit and at rest, and access control mechanisms including multi-factor authentication and role-based access.
4. Privacy
How transparent, consistent, and user-centric is the product’s data handling? We assess privacy policy clarity and enforceability, compliance certifications (SOC 2, GDPR, CCPA/CPRA, HIPAA where applicable), third-party data sharing practices, data retention and deletion controls, and user rights around access, portability, and consent.
5. Sustainability & Reliability
Evaluated across two dimensions: company sustainability — funding status and runway, revenue trajectory, team depth, and market position — and service reliability — uptime commitments, published SLAs, historical incident records, and performance benchmarks. Environmental and social practices are noted where publicly documented.
6. Impact, Ethics & Safety
Does the product and its developer act responsibly toward users and society? We examine the existence and substance of a published ethics policy, bias mitigation practices, model transparency, controls to prevent misuse and harm, societal applications and misuse risk, and the company’s track record of responding to ethical failures or safety incidents.
We Run Our Own Tests
huby doesn’t evaluate a product only from what others have published about it. We generate our own primary evidence, and it sits at the top of our evidence hierarchy.
When we take on a category of AI products, we identify a cohort of products that compete in it. For each of our six evaluation categories we establish a common set of subcategories and atomic factors, and a common set of test criteria. Then we run those tests ourselves.
Every test is run identically across every product in the cohort — the same inputs, within as narrow a time window as we can manage. That consistency is what makes a comparison meaningful. A test run against one product in January and a rival in June is measuring two different moments, not two different products. We test whatever access is available to us: the publicly available free or trial tier, or a test account or API key provided by the product owner.
Scoring: Evidence-Anchored, 1.0 to 5.0
All scores are reported on a 1.0–5.0 scale at atomic factor, subcategory, and category levels. Each band has a specific evidence requirement — scores are not assigned by feel.
| 5.0 | ExemplaryStrong, independently verified evidence of best-in-class practice. Formal audits, third-party certifications, or peer-reviewed assessments confirm performance. |
| 4.0–4.9 | StrongMultiple credible, corroborated sources substantiate the claim. Minor gaps exist but don’t materially undermine the finding. |
| 3.0–3.9 | AdequateEvidence is present but limited, mixed, or partially corroborated. Basic standard met, with notable gaps or inconsistencies. |
| 2.0–2.9 | WeakMeaningful shortfalls. Claims made by the product owner are not substantiated by independent sources, or documented failures exist. |
| 1.0–1.9 | DeficientMaterial failures or absence of the capability or control. Significant risk to users or enterprise buyers. |
Weighted Aggregation — Locked Before Data Collection
Scores aggregate upward through the three-tier hierarchy using a weighted model established during framework design — not a simple average, and not a model that can be adjusted once we’ve seen the results.
Within each subcategory, atomic factors carry defined weights reflecting relative criticality. For example, the presence of multi-factor authentication in an Access Control subcategory is weighted more heavily than audit log granularity, because its absence represents more material risk. Weights within each subcategory sum to 100%. The same logic applies from subcategory to category, and category scores roll up in turn into a single composite score for the product. You can see the assigned score at every level, not just the headline number.
A Five-Tier Source Hierarchy
Not all evidence is equal. huby uses a tiered source framework where higher-tier sources take precedence when sources conflict.
- Tier 1
Primary research and testing by huby
Identical tests run across every product in a cohort, against the framework for that category; test date, product version, access tier and sample size recorded - Tier 2
Independent third-party audits and certifications
SOC 2 reports, CVE databases, named security firm assessments, peer-reviewed publicly available research - Tier 3
Product owner documentation
Official privacy policies, terms of service, technical documentation, changelogs - Tier 4
Structured user and practitioner evidence
Enterprise case studies from named customers, verified professional community posts - Tier 5
Anecdotal community evidence
Reddit threads, social media posts, informal reviews
Every claim that affects a score has to be traceable. Sources in Tiers 2 through 5 are cited with a publicly accessible URL. Tier 1 findings are our own measurements and have no external URL to point at, so they are documented differently: the test date, the product version and access tier tested, the sample size, and a link to the published test methodology — so that a reader can reproduce the test rather than simply follow a citation.
What Independence Actually Means Here
huby does not charge product owners for evaluation, for listing, or for either report. There is no paid tier, no expedited review, and no way for a product owner to pay for placement, for a score, or for a score to be reconsidered.
Evaluations are not contingent on a product owner engaging with us at all — most are initiated by huby without the owner’s involvement. For newer products without a long public evidence trail, owners may submit documentation directly; those submissions are identified as such in citations and receive no preferential weighting. Should any commercial relationship with a product owner exist in future, it will be disclosed in the relevant report, and the methodology will say so first.
Evidence Gaps Are Disclosed, Not Buried
Where a product capability cannot be assessed due to insufficient public evidence, we record it explicitly as an evidence gap — we don’t infer from absence, and we don’t assign a low score based on speculation. This is disclosed in the report and may result in an N/A score for that factor rather than a low one. Where sufficient public data exists, scores are also contextualised against the product’s competitive set.
Two Report Types for Different Audiences
Public Transparency Report
Covers all six evaluation categories with subcategory-level scores and written assessments. Evidence is cited with verifiable URLs. Available to all huby users at no cost.
Detailed Owner Report
Produced for the product owner. Includes atomic factor-level scores, specific improvement recommendations mapped to each factor, and competitive gap analysis. Confidential to the owner and provided free of charge — receiving it doesn’t require having submitted the product, and doesn’t affect any score.
Each report carries a production date and is reviewed for material updates on a regular basis. A revision can be brought forward when we’re notified of a significant product change, a security incident, or a regulatory development that warrants earlier review. Score changes between versions are tracked and published with explanatory notes, so you can see what moved and why.
Limitations Worth Knowing
huby reports are grounded in our own testing, publicly available evidence, and product owner submissions at the time of evaluation. They do not constitute a formal security audit, legal compliance certification, or financial advisory opinion. Scores reflect the state of the product at the report production date and may not reflect subsequent changes.
Our own tests are conducted on a defined sample of test items, on a specific product version and access tier, at a point in time. They are not exhaustive. Sample size and test scope are disclosed in each report, and results drawn from small samples carry correspondingly wide uncertainty — they indicate direction rather than a precise ranking. Where testing is performed on a free or trial tier, the results describe that tier; paid tiers may use different models or configurations and may perform differently.
Importantly, the absence of evidence for a capability or control does not confirm its absence — it reflects the limits of publicly available information at evaluation time. And we are not lawyers, certified security auditors, or financial advisors. Where legal, security, or financial interpretation is required, we say so and encourage readers to seek qualified professional advice.
Why This Approach?
AI product evaluation is genuinely difficult. Evidence is scattered, vendor claims are optimistic, and the stakes for enterprise buyers making adoption decisions are real — security exposure, regulatory risk, wasted investment.
We built our methodology to be transparent precisely because we believe the credibility of our findings depends on readers being able to assess the validity — and the limitations — of our work for themselves. The framework is public. The weights are published before data collection begins. The gaps are disclosed rather than papered over.
That is what independent evaluation actually means.
Ready to see the methodology applied to a real product? Browse our AI product evaluations on huby.