News

Vals Aims to Set the Standard for AI Model Evaluation

Vals raised a $40 million Series A led by Andreessen Horowitz and focuses on private evaluations for complex industry tasks and federal agencies.
Sep 20, 20263 min read
Vals Aims to Set the Standard for AI Model Evaluation

#

Key takeaways

  • Vals raised a $40 million Series A led by Andreessen Horowitz.
  • The company evaluates AI models on complex industry tasks, not only public tests.
  • Its private evaluations aim to measure practical performance and negative outcomes.
  • Vals has launched an evaluation program for federal agencies.

Vals raises funding and sharpens its focus

TechCrunch reports that Vals has raised a $40 million Series A led by Andreessen Horowitz. The startup is positioning itself around a specific problem in AI evaluation. It wants to become the gold standard for how models are judged.

The company does not focus only on public tests. Instead, it evaluates AI models on complex industry tasks. That approach suggests a broader view of model quality. It also points to a stronger emphasis on real-world usefulness.

What Vals says it measures

Vals says its private evaluations are meant to measure practical performance. They also aim to measure negative outcomes. That distinction matters because a model can look strong in a public benchmark and still fail in a more realistic setting.

The source does not provide technical details about the evaluation method. It also does not describe the exact industries covered. Based on the report, the main idea is clear: Vals is trying to test models in ways that reflect actual work more closely.

Why private evaluations matter

Public tests can be useful, but they may not capture every risk or limitation. Vals is betting that private evaluations can reveal more about how a model behaves in practice. That includes both performance and harmful side effects.

This approach may appeal to organizations that want more than a score on a leaderboard. It may also help buyers compare models on tasks that resemble their own workflows. That is an assumption based on the company's stated focus, not a claim from the source.

Federal agency program

Vals has also launched an evaluation program for federal agencies. The source does not explain how the program works. It does not say which agencies are involved or what tasks they will test.

Even so, the launch suggests that Vals wants its evaluation model to be used in formal settings. That could make the company's work relevant to public-sector decision-making. The report does not provide enough detail to go further.

What this means for AI buyers

The report highlights a growing interest in evaluation methods that go beyond public benchmarks. Buyers may want to know how a model performs on specific tasks, not just general tests. They may also want to understand failure modes before deployment.

Vals appears to be building around that need. Its pitch is not about making models. It is about judging them more carefully. That can be valuable when performance and risk both matter.

Operational and governance considerations

The source points to two practical concerns: real-world performance and negative outcomes. Those are both important when choosing or reviewing AI systems. A model that performs well in one setting may behave differently in another.

Private evaluations can help surface those differences. They can also support more disciplined procurement and oversight. The report does not describe a governance framework, so any broader process would be an assumption.

Morocco relevance

The source reports no Morocco-specific facts. For readers in any market, the global lesson is simple: evaluate AI systems on tasks that resemble real use, not only on public tests.

Bottom line

Vals is trying to make AI evaluation more practical. Its funding round gives it more room to pursue that goal. The company's focus on private testing, negative outcomes, and federal agency use cases sets it apart from benchmark-only approaches.

The report does not prove that this method is the best one. It does show that evaluation is becoming a more serious part of the AI stack. For many buyers, that may matter as much as model performance itself.

Follow us on Google

Add Intelligence Artificielle Maroc as a preferred source to see more of our relevant stories in Google Search.

Add us as a preferred source
AI platform development

What would you like to build?

We build custom AI platforms, SaaS products, intelligent business applications, and automation systems.

This form is for project inquiries, not general questions about artificial intelligence.

Name *
Work email *
Organization (optional)
Solution *
Short project description *

Related Articles

featured
J
Jawad
Sep 20, 2026

TypeSafe AI's Jev model focuses on calibrated decisions

featured
J
Jawad
Sep 20, 2026

Vantora raises $100M to build physical-AI startups

featured
J
Jawad
Sep 20, 2026

AI hallucination nearly triggers US military operation

featured
J
Jawad
Sep 20, 2026

Anthropic says Claude now writes most merged code internally