Field Journal · Issue 1029

Why Cosmos AI Lab Wants to Build a New Standard for Measuring Human AI Capability

As AI systems outperform humans in knowledge and precision work, traditional credentials are losing their power as signals of competence. Cosmos AI Lab is building a new framework to measure human strategic capability when working with AI — not what you know alone, but what you can achieve when integrated with non-biological intelligence.

Alex Pham · · 8 min read · Updated June 23, 2026
#engineering
Read in
Why Cosmos AI Lab Wants to Build a New Standard for Measuring Human AI Capability

If AI and robotics are steadily eroding the traditional advantages of human expertise — and I've argued elsewhere that they are — then a new question immediately presents itself: how will society evaluate human competence in the AI era?

For centuries, the answer has been straightforward. Degrees, certifications, years of experience, professional titles, employment history. These signals worked because they reflected a world where knowledge was scarce, execution was limited by individual ability, and human specialization was the core asset of any economy.

That world is disappearing.


The credential gap

Here's something that anyone working seriously with AI has already noticed: credentials have become poor predictors of actual output quality.

A person with a modest educational background who knows how to select the right AI for the right task, decompose a project into manageable modules, design verification checkpoints, and coordinate multiple AI systems into a coherent workflow can now produce results that surpass what a highly credentialed professional achieves using AI poorly — or not at all.

This isn't hypothetical. It's happening right now, across industries, and the gap is widening fast.

What this reveals is that the labor market is developing a need for entirely new competence signals — signals that don't measure "what a person knows standing alone," but rather "what a person can accomplish when integrated with their AI systems." And right now, no credible, standardized framework exists to provide those signals.

That's the gap Cosmos AI Lab is trying to fill.


Why traditional testing fails here

The moment you mention "measuring AI capability," most people instinctively think of quizzes, multiple-choice tests, essay prompts, or timed coding challenges. The familiar apparatus of assessment.

But this is exactly where old thinking breaks down.

In any open online assessment, participants will — obviously — use AI to answer questions. If the test still relies on questions with clear correct answers, then the results don't actually measure human capability anymore. They primarily measure which AI the person has access to, how well they prompt it, and whether they're paying for a premium tier. That's not a competence assessment. That's a tool-access survey.

A meaningful evaluation system for the AI era has to abandon the logic of "what do you know?" and replace it with something fundamentally different: "how far can you go in a complex environment when you're allowed to use every AI and tool at your disposal?"

That's not a minor adjustment to existing testing. It's a complete philosophical shift in what assessment means.


How Cosmos AI Lab thinks about this problem

We don't see this as building "an AI quiz." We see it as designing AI-native capability measurement — evaluation built from the ground up for a world where AI is always present.

The key principles are simple but radical in their implications. Don't ban AI. Don't try to separate humans from their tools. Don't design assessments around "catching cheaters" by forcing people to work unaided. Instead, assume AI is always there. Make it the default.

Once you accept that, the question shifts completely. It's no longer "can AI answer this question?" — of course it can. The real question becomes: "how effectively can this human leverage AI to push a genuinely difficult problem closer to a solution?"

That's why we're interested in evaluation models like open-world trials where participants face realistic, unstructured challenges; simulated environments that mirror real-world complexity; frontier tasks deliberately set at the edge of what current AI can handle; outcome-based scoring that judges actual results rather than process compliance; progress-based evaluation that measures how far someone advances rather than whether they reach a final answer; and orchestration benchmarks that specifically test the ability to coordinate multiple AI systems toward a coherent goal.


Measuring progress, not just success

One of the deepest flaws in conventional assessment is its binary nature: pass or fail, correct or incorrect, solved or unsolved.

But many of the most important problems humanity will face in the coming decades don't have clean solutions that someone can simply arrive at. In complex system optimization, scientific simulation, novel AI design, and multi-variable open environments, even the best human-AI combinations may not fully solve the problem. Yet the differences between individuals remain enormous in terms of how far they get, how intelligently they approach the challenge, and how much they narrow the gap between the starting point and a viable solution.

A person who makes no progress is fundamentally different from a person who advances 60% toward a solution but can't close the final gap. Current assessment systems treat both the same: "didn't solve it." That's absurd, and it throws away exactly the signal that matters most.

This is why Cosmos AI Lab focuses on what we call trajectory quality — not just whether someone succeeded, but how they moved through the problem space. The quality of their strategy, their efficiency in deploying AI resources, their ability to decompose problems, their skill in cross-validating outputs, their adaptability when conditions change, and their capacity to manage competing objectives within constraints.

These dimensions tell you far more about someone's real capability than any binary pass/fail ever could.


What we're actually trying to measure

Human AI capability isn't a single skill. It's a composite of several distinct competencies that interact with each other.

AI Literacy — understanding the strengths and limitations of different model types, recognizing hallucination patterns, working within context window constraints, and managing privacy and risk considerations.

Prompt and Instruction Design — the ability to define clear objectives, describe requirements precisely, specify output formats, and provide exactly enough context for the AI to perform well without being overwhelmed or misdirected.

AI Orchestration — knowing how to break a problem into components, assign the right AI to each component, design review and checking mechanisms, and prevent conflicts between different modules working on the same project.

Workflow and System Building — integrating AI into real processes, connecting with tools, APIs, and databases, and constructing guardrails that keep the system stable under real-world conditions.

Judgment and Verification — evaluating AI output critically, detecting errors, assessing confidence levels, and knowing when to escalate to human review or specialized expertise.

Strategic Endurance — maintaining coherent strategy across tasks that span hours or days, where no quick solution exists and continuous adaptation is required.

Taken together, these form what we call human strategic AI utilization capacity — the strategic ability of a human to use, orchestrate, and extract value from AI systems. It's this composite capability, not any single skill within it, that determines who can actually create value in an AI-native environment.


Why this could become the credential of the future

In the emerging economy, organizations won't just need people "with degrees." They'll need people who can work fluently across multiple AI systems, translate AI capability into concrete outcomes, manage the risks that come with AI-dependent workflows, operate effectively in rapidly changing environments, and build systems that scale.

If a measurement framework can reliably, objectively, and predictably assess those capabilities — and if that assessment correlates with real-world job performance — then it becomes something genuinely powerful: a new type of credential.

Not because it erases traditional degrees. But because in many contexts, it will be a more accurate signal of actual productive capacity. Companies making hiring decisions, investors evaluating teams, research organizations building new capabilities, even educational institutions redesigning curricula — all of them will need data that reflects this new reality.

The question they'll need answered is straightforward: who can actually operate in an AI-native environment at a strategic level? Traditional credentials don't answer that. Something new has to.


Infrastructure, not a quiz

To be clear about the ambition: Cosmos AI Lab isn't building a test. We're building evaluation infrastructure.

That means open simulation environments where real complexity is preserved rather than stripped away. Frontier tasks calibrated at the boundary of current AI capability so they can't be trivially solved by any single model. Short-form assessments that take two hours alongside operational tests spanning 24 hours and campaign-length evaluations running over seven days. Scoring systems that weigh outcomes, efficiency, endurance, and proximity to solution. Dynamic capability profiles that evolve across domains as a person's skills develop. And evaluation standards specifically designed so that no single powerful model or clever prompt can game the system.

If built correctly, this isn't just an assessment product. It's a new signal layer for the labor market, a new reference framework for competence in the AI era, and potentially a foundational piece of infrastructure for the AI-native economy that's taking shape whether we're ready for it or not.


The question that matters

In the old world, credentials answered: "What have you learned?"

In the new world, the question that matters is different: "What can you accomplish with the non-biological intelligence you command, orchestrate, and build?"

That's the gap Cosmos AI Lab exists to fill. We believe that skills like AI utilization, multi-AI orchestration, AI workflow design, output verification, and strategic operation in AI-rich environments will become core human competencies — not because they're trendy, but because they may be the new foundation for determining who can genuinely create value in a world where traditional expertise is being fundamentally restructured.

And if society truly needs an objective, verifiable, scalable, and widely recognized way to measure those competencies, then building that framework won't remain an interesting option.

It will become a necessity.


This article is part of a series exploring AI's impact on human expertise, credentials, and social value. Published by Cosmos AI Lab — a crypto and AI research startup based in Vietnam.

Share this issue