Demonstrations
Worked examples from people who do the job: clinical reasoning, contract analysis, financial modelling, production code.
Two kinds of people are scarce right now. Specialists who can demonstrate, grade and break a model in their own field. And engineers who will sit inside your team and get the thing into production. We supply both, as one deployed unit.
Training · evals · forward-deployed engineering
Labs and enterprises are well served at both ends. There are vendors who will label at volume, and there are consultancies who will write you a strategy. The gap is the part where a model has to become something that works on a Tuesday, in your systems, against your data.

Expert demonstrations, preference ranking, evaluations and red-teaming, produced by people who hold the credential the task needs.

Forward-deployed engineers embedded in your team, closing the distance between a model that scores well and a system your business actually runs on.

The agents, tools and evaluation harnesses that turn a capability into a product. We build these for ourselves, which is why the engineers are worth deploying.
Commodity labelling stopped being the bottleneck. What is scarce is a specialist who understands the task well enough to demonstrate it, grade it, or break it.
Worked examples from people who do the job: clinical reasoning, contract analysis, financial modelling, production code.
Pairwise judgments against your rubric. Inter-rater agreement gets measured and reported with every batch.
Rubric design, benchmark grading and regression suites, built by people who run eval harnesses on their own systems.
Adversarial probing and jailbreak discovery, written up so your safety team can reproduce and act on every finding.
Native fluency in Urdu, Punjabi, Sindhi and Pashto alongside professional English — coverage most vendors quietly outsource.
Task-based testing in real environments — browsers, codebases, workflows — scored against criteria you define up front.
You brief a pod once. It keeps the context, the corrections and the standard, and it is still there next quarter.
Graduates and practising professionals out of LUMS, NUST, IBA, FAST and GIKI, plus licensed specialists where the work needs a credential.
Language assessment, domain testing and a reasoning screen. We publish the pass rates so you can see what the filter actually removes.
They train against your guidelines until agreement clears the threshold you set, and we absorb the cost of getting them there.
Every batch arrives with sampling results, double-marked disagreements and the agreement rate attached.
Context compounds instead of resetting. Your rubric does not get re-explained to a new stranger every week.
An FDE sits with the people who own the problem, writes production code against your stack, and stays long enough to be held to the outcome. Ours come off systems we operate every day, so the first week is spent on your problem rather than on learning what an agent is.
In your standups, your repo and your incident channel. The distance between the person who understands the workflow and the person writing the code goes to zero.
Agents, tools, retrieval, evaluation harnesses and the plumbing between them. Reviewed by your engineers, merged into your repositories, owned by you.
A capability nobody can measure is a capability nobody will sign off. The harness that proves the thing works gets built alongside the thing itself.
The engagement ends with your team running the system without us. Anything else is a dependency we sold you, and it would show up in the second invoice instead of the first.
The same labs paying US rates for expert hours are competing for people who work remotely anyway. Pakistan already exports this work through freelance platforms at scale. We package it with vetting, calibration and accountability.
Procurement asks the same questions every time. Here are the answers before you ask: every contributor signs an NDA before their first task, access is least-privilege and revoked on rotation, and work runs inside your stack when your policy requires it.
SOC 2 Type II and ISO 27001 are still on the roadmap. Until they land we start you on non-sensitive and public-data pilots, and we say so up front rather than letting you find out in diligence.
Signed by every contributor. Client work is never reused, resold or shown in a portfolio.
Scoped credentials, revoked on rotation, with an audit trail per batch.
Sampling, double-marking and agreement rates reported with the work.
A bounded batch before any commitment. Judge the output, then scale the pod.
We work inside your annotation stack, so nothing migrates and no data leaves your perimeter.
A pod lead who answers for quality, reachable in your timezone overlap.
Most teams have already tried at least one of these. The differences show up in rework, not in the first invoice.
| Crowd marketplace | Traditional BPO | Kitsune pod | |
|---|---|---|---|
| Who does the task | Whoever claims it that hour | Whoever is on shift | A named pod that stays on your account |
| Domain credential | Self-declared | General graduate pool | Tested, and licensed where the work needs it |
| Calibration to your rubric | You write guidelines and hope | Billed to you as training hours | Done before billing starts, at our cost |
| Quality evidence | Spot checks you run yourself | A throughput report | Agreement rate and double-marked disagreements per batch |
| Context between batches | Resets every time | Resets on rotation | Same people, corrections compound |
| Who answers when quality slips | Platform support | An account manager | The pod lead who signed the batch |
| Language coverage | English-first | English-first | Professional English plus Urdu, Punjabi, Sindhi, Pashto |

You send the task, the volume and the standard. We come back with the sample design, the pass criteria and what the batch will cost, in writing.

Contributors are screened for your domain and trained against your guidelines until agreement clears the threshold. Those hours are ours, not yours.

Work lands with sampling results, the agreement rate and every double-marked disagreement attached, so your reviewers can audit it instead of trusting it.
A single bounded pilot batch. We would rather you judge a small piece of real output than sign a volume commitment against a capability deck. If the pilot does not clear your bar, that is the end of it and you keep the work.
Language assessment, a domain test written against the task you are actually buying, and a reasoning screen. Where the work needs a licence — clinical, legal, financial — we verify the credential instead of taking a claim on a CV.
A sample of every batch is double-marked by a second qualified contributor. We report the inter-rater agreement rate and hand over the disagreements themselves, so you can see where the rubric is ambiguous instead of only where the workers were wrong.
Not yet. Both are on the roadmap and neither has landed. Until they do we start clients on non-sensitive or public-data work, and we say so here rather than letting it surface halfway through your vendor review.
Yes, and it usually should. We work inside your annotation stack under scoped credentials, so your data stays in your perimeter and nothing has to be migrated to us and migrated back.
You do, on delivery. Client work is never reused for another engagement, resold, or shown in a portfolio, and every contributor signs an NDA before their first task.
The country already exports this class of work at scale through freelance platforms, at rates well under US and European equivalents, with a large English-fluent graduate pipeline out of LUMS, NUST, IBA, FAST and GIKI. What has been missing is a vendor wrapping that supply in vetting, calibration and accountability. That is the whole business.
Pakistan Standard Time is UTC+5, which gives you a full working overlap with Europe, the Gulf and India, and a morning overlap with the US east coast. Pod leads are reachable inside your working hours, not at the end of a queue.
A calibrated pod grows faster than a cold one because the standard is already documented and the existing members train the new ones. Realistically, doubling a pod takes two to three weeks including screening and calibration, and we will tell you if your timeline needs more than that.
Tell us the task, the volume and the standard. We will scope a bounded pilot, run it with a calibrated pod, and hand it back with the quality evidence attached.