What we do

Narrative & PR Get published in the outlets that matter Agentic Systems AI that runs marketing and media Creator OS The business layer for creators OdinVision Content operations on autopilot YouTube Network Claims, reuse and geo revenue Frontier Data Human data for frontier models

Company

Our vision Built for humans Founder The Kitsune AI story Life & people Our values & advisors

Resources

Our Work What we built & the resultsBlog Honest writing from the team Press & media The official record
Let's talk AI
Frontier Data

The people who make your model worth shipping.

Doctors, lawyers, engineers, analysts and linguists — screened, calibrated to your rubric, and held to an agreement rate you can audit. Expert demonstrations, preference data, evaluations and red-team findings, delivered on your schedule.

Specialists working on model evaluation Vetted pods · calibrated QA
What we produce

The data your team keeps queuing for.

Commodity labelling stopped being the bottleneck. What is scarce is a specialist who understands the task well enough to demonstrate it, grade it, or break it.

Domain specialists writing demonstrations
SFT · expert generation

Demonstrations

Worked examples from people who do the job: clinical reasoning, contract analysis, financial modelling, production code.

Reviewers comparing model responses
Preference · RLHF

Ranking and preference

Pairwise judgments against your rubric. Inter-rater agreement gets measured and reported with every batch.

Evaluation harness being run
Evals

Evaluation and grading

Rubric design, benchmark grading and regression suites, built by people who run eval harnesses on their own systems.

Adversarial testing session
Safety

Red-teaming

Adversarial probing and jailbreak discovery, written up so your safety team can reproduce and act on every finding.

Multilingual reviewers at work
Multilingual

Urdu and regional languages

Native fluency in Urdu, Punjabi, Sindhi and Pashto alongside professional English — coverage most vendors quietly outsource.

Agent task evaluation
Agents

Agent evaluation

Task-based testing in real environments — browsers, codebases, workflows — scored against criteria you define up front.

How you get it

One team that stays on your account.

You brief a pod once. It keeps the context, the corrections and the standard, and it is still there next quarter.

Pod lead reviewing work with specialists
  1. 01

    We source for your domain

    Graduates and practising professionals out of LUMS, NUST, IBA, FAST and GIKI, plus licensed specialists where the work needs a credential.

  2. 02

    They get screened before they touch your data

    Language assessment, domain testing and a reasoning screen. We publish the pass rates so you can see what the filter actually removes.

  3. 03

    The pod calibrates to your rubric

    They train against your guidelines until agreement clears the threshold you set, and we absorb the cost of getting them there.

  4. 04

    Work ships with its own evidence

    Every batch arrives with sampling results, double-marked disagreements and the agreement rate attached.

  5. 05

    The same people stay on your account

    Context compounds instead of resetting. Your rubric does not get re-explained to a new stranger every week.

Where the supply comes from

Pakistan is the talent pool the industry has not priced in.

The same labs paying US rates for expert hours are competing for people who work remotely anyway. Pakistan already exports this work through freelance platforms at scale. We package it with vetting, calibration and accountability.

$4.5BPakistan IT exports, FY2026, up 29% year on year
$950MFreelance earnings in ten months
Top 5Country on global freelance platforms
5Universities feeding the vetting funnel
Supervised delivery floor

Built to pass your vendor review.

Procurement asks the same questions every time. Here are the answers before you ask: every contributor signs an NDA before their first task, access is least-privilege and revoked on rotation, and work runs inside your stack when your policy requires it.

SOC 2 Type II and ISO 27001 are still on the roadmap. Until they land we start you on non-sensitive and public-data pilots, and we say so up front rather than letting you find out in diligence.

NDA before first task

Signed by every contributor. Client work is never reused, resold or shown in a portfolio.

Least-privilege access

Scoped credentials, revoked on rotation, with an audit trail per batch.

Measured QA

Sampling, double-marking and agreement rates reported with the work.

Pilot first

A bounded batch before any commitment. Judge the output, then scale the pod.

Your tools

We work inside your annotation stack, so nothing migrates and no data leaves your perimeter.

Named supervision

A pod lead who answers for quality, reachable in your timezone overlap.

Start small

Send one batch. Judge the output.

Tell us the task, the volume and the standard. We will scope a bounded pilot, run it with a calibrated pod, and hand it back with the quality evidence attached.