A5 / Incentive design
Can bounties and rewards make AI-assisted teams move faster?
Sometimes: rewards can speed a real queue when they pay only for verified outcomes, but weak metrics can increase activity without improving useful work.
- 01Prepare the brief
- 02Pair human + agent
- 03Verify the outcome
Sometimes—but points are the least important part.
A leaderboard can make a dull dashboard brighter without changing the work. A real incentive system does something harder: it names an outcome the organization values, requires evidence that the outcome happened, puts an independent decision between the claim and the reward, and sends scarce resources back toward useful work.
That can redirect attention. It can also redirect attention toward the wrong thing. The difference is not whether the interface looks like a game. The difference is what the system rewards, what it refuses to count, and what happens when the metric starts winning over the mission.
Key takeaways
- Rewards can redirect attention and resources, but observed activity after a rule change does not by itself prove that the reward caused the outcome.
- A defensible bounty loop requires a valued outcome, evidence, independent acceptance, and active monitoring for gaming, neglected work, and unequal access.
- Contribution data can inform a reviewed, time-bounded tool allocation; it should not automatically decide who receives the best AI subscription.
A wartime case—and where the comparison ends
Ukraine's Army of Drones Bonus program is an unusually visible example of a verified-outcome reward loop. It is also a response to Russia's invasion, built around combat, loss, and national survival. Office work is not war. Completing a software task is not morally or practically comparable to a battlefield strike. The lesson for a company is limited to the design of a feedback system: evidence, verification, priorities, allocation, and the risks created by every proxy.
The program began in August 2024, according to TIME's reporting from Ukraine. Its procurement loop was expanded the following year by connecting three systems:
- units upload photo or video evidence into the secure DELTA environment;
- verified outcomes produce e-points shown in Brave1 Market;
- authorized unit personnel use those points to order additional equipment, while DOT-Chain Defence handles contracts, logistics, and delivery records.
That is the process described in the Ukrainian Ministry of Defence's September 2025 guide. The ministry is explicit that these orders supplement centralized supply rather than replace it. In other words, the reward is added capacity, not the withdrawal of a basic entitlement.
By January 2026, the ministry said units had ordered 240,000 drones through Brave1 Market over six months and that more than 160,000 drones and other items had been delivered. It also reported 819,737 video-confirmed strikes during 2025. Those numbers show that the recording and procurement loop reached substantial scale. They do not show how many outcomes the points caused. The official reports do not provide a control group or a counterfactual for what the same units would have achieved under different supply, operational, or reporting conditions. See the ministry's marketplace update and 2025 program report.
This distinction matters because the boldest causal claim came from a before-and-after comparison. Mykhailo Fedorov told TIME that a change in the points assigned to one type of target was followed by a doubling of the recorded count in a month. That is evidence that behavior or reporting moved after the rule changed. It is not, on its own, proof that the rule produced the whole increase: participation, battlefield conditions, equipment availability, targeting priorities, and verification volume could all have moved at the same time.
The case also exposes the failure modes instead of hiding them. Reporting in TIME described commanders who initially rejected the program and warned that concentrating performance data or publicizing leading units created security risks. A later Guardian report noted that commanders still had to follow operational priorities when the correct mission carried fewer points. Ukrainska Pravda reported a concrete blind spot: casualty-evacuation work was not receiving e-points at the time of its report. By January 2026, however, Ukraine's Ministry of Defence said verified robotic evacuation missions were receiving e-points. A 2026 analysis in International Politics raises broader strategic and ethical concerns about rewarding a measurable tactical proxy.
So the defensible lesson is not “gamification made Ukraine more effective.” It is narrower and more useful: a closed loop can turn verified field outcomes into faster, more decentralized allocation—but its designers must keep correcting what the score cannot see.
The transferable mechanism is a loop, not a leaderboard
Strip away the theme and the reusable pattern looks like this:
valued outcome -> claim -> evidence -> independent verification
^ |
| v
priority review <- measured side effects <- credit -> approved resource
Each arrow matters.
The outcome must be worth having. “Used more tokens,” “sent more prompts,” and “closed more tickets” are easy to count. None proves that useful work reached the company. The target should be an accepted result: a fixed defect, a tested asset, a finished investigation, or a decision made with adequate evidence.
Evidence must travel with the result. A claim becomes auditable when it links to the test, recording, artifact, branch, source, or observation that supports it. Evidence does not eliminate judgment. It gives the reviewer something better than confidence to judge.
Verification must be meaningfully independent. If the same person defines the reward, claims the work, declares success, and pays themselves, the loop invites score creation rather than value creation.
The reward should improve the next useful action. Ukraine's loop returns equipment to units. A company can return scarce AI capacity, protected focus time, training, recognition, or a bounded bonus. The safest reward is often capability, not cash.
The base system cannot disappear. Bonus allocation should not deprive people of the tools, pay, access, or psychological safety required to do their jobs. A bonus is a steering mechanism around the margin, not a substitute for fair provisioning.
What Wagglet implements today
Wagglet's current loop is intentionally smaller than a redeemable economy.
Every task has an owner-selected Work Value. The five levels show runner awards of 1, 3, 9, 15, or 20 Contribution Score points after acceptance. The scale gives a task a visible value without pretending that hours, difficulty, urgency, and impact are the same number.
If a published task remains open and unclaimed, one of its owners can add a bounty in steps from +1 to +5, up to +25 total. The bounty is extra runner score for a stuck task. It does not increase the task owner's award. When somebody claims the task, that bounty is frozen on the attempt so the promised amount cannot change underneath the runner.
The score arrives only after a delivery is accepted. An independently accepted close pays the runner's Work Value plus the frozen bounty. A self-verified close receives a reduced award and pays no bounty. The task's attempts, delivery evidence, outcome, final value, recipient, and award remain explainable through durable records. Task creators can also receive credit when another person completes useful work they defined, so writing a good task is not treated as invisible management overhead.
Consider a task with runner value 9 that has sat untouched long enough for its owner to add a +4 bounty:
9 runner value + 4 stuck-task bounty = 13 after independent acceptance
Claiming it earns zero. Running an agent earns zero. Producing a delivery earns zero until it is reviewed. If the delivery is sent back, the attempt remains part of the quality history but the reward does not appear. The number is attached to a verdict on an outcome, not to activity around it.
Wagglet also keeps an all-time Contribution Score and exposes delivery, first-pass, and rejection-related facts in its statistics surfaces. But the score is deliberately non-spendable. It is not money, a wallet, or a token that Wagglet currently exchanges for a Claude Code or Codex subscription. Any “earn a better subscription” policy is a future company policy or product decision—not a feature available today.
Design for the employee who understands the metric perfectly
The most useful incentive-design question is not “Will people cheat?” It is: What happens when a smart, well-intentioned person optimizes exactly what we asked for?
Research gives good reason to ask. An experiment on points, levels, and leaderboards found more output quantity without a corresponding improvement in intrinsic motivation or competence. A field study of explicit performance incentives found strategic reporting that damaged the organization's real objective. Different settings will produce different effects, but together they warn against treating a visible count as the work itself. See the studies in Computers in Human Behavior and the Journal of Labor Economics.
A resilient AI-work reward system therefore needs at least these controls:
- Reward accepted outcomes, never token consumption. Burning a subscription quota is a cost signal, not value. Prompt count, agent time, and model spend are diagnostic inputs at most.
- Require task-specific proof. “Tests pass” may be enough for a bounded code change; a visual feature may need human observation; a research task needs sources and an uncertainty statement. One universal proof field will be gamed because it fits nothing well.
- Keep an independent acceptance path. Self-verified work can remain possible, but it should be visibly different and should not unlock the same bounty.
- Measure side effects beside the score. Review rejections, reopened work, first-pass outcomes, queue age, time to delivery, and downstream defects. A rising points total with rising rework is not a win.
- Bound every multiplier. Small steps and a hard ceiling limit damage while a team learns. Wagglet's +25 bounty cap is one example.
- Audit opportunity, not only results. People in visible technical roles may have more chances to earn points than office managers, artists, reviewers, or those doing preventive work. Compare access to claimable work across roles and accommodate disability, caregiving, time zones, and different working patterns.
- Keep judgment and appeals. A reviewer must be able to reject weak proof; a manager must be able to override a mechanically bad allocation; an employee must be able to contest a mistaken outcome without bargaining over a public leaderboard.
- Do not turn the scoreboard into surveillance. Collect the least evidence needed to verify the task, restrict sensitive details, and define retention before rewards make over-collection feel useful.
This is also why team competition is often safer than a permanent individual race. Individual ranks can discourage help, task writing, review, and honest escalation—the work that makes another person's accepted outcome possible. A cooperative target with role-specific credit preserves more of the system around the visible result.
Could Contribution Score help allocate better AI subscriptions?
Yes, as one input to a reviewed company policy. No, as an automatic vending machine.
A sensible pilot might let an employee or team request a higher Claude Code or Codex plan after capacity limits repeatedly block valuable accepted work. The request could link to recent outcomes, documented quota waits, and the upcoming task queue. A manager would then approve a time-bounded upgrade from a defined budget and review whether it reduced delays.
That policy should never mean “the top scorer gets the best tools.” Different roles have different opportunity to earn score, and historical output is not identical to future need. Nor should teammates share provider accounts or credentials: any allocation must follow the company's security policy and the provider's current plan and licensing terms.
The clean formulation is:
verified contribution + demonstrated capacity need + suitable upcoming work
-> manager-reviewed trial allocation
Contribution is evidence. Capacity need is the reason. Human review is the fairness and budget boundary.
Run a bounded experiment before announcing a revolution
For a first Wagglet pilot, choose one queue where work genuinely sits unclaimed. Do not gamify the entire company.
Record a baseline for queue age, time from claim to delivery, acceptance rate, rework, and who gets access to the work. For 30 days, let task owners add small bounties only after an agreed waiting period. Require a concrete delivery report and independent acceptance. Review the data weekly, including tasks people avoided and important work that earned no points.
Stop or change the pilot if low-value tasks multiply, reviewers become a bottleneck, people hoard work, quality falls, or access to rewards concentrates by role. Continue only if the queue moves without those costs—and describe the result as evidence from that team and time window, not a universal law.
That is the serious promise of gamification for AI-assisted work. Not confetti. Not “engagement.” A visible, bounded agreement about which outcomes matter, what proves they happened, and how useful contribution earns the capacity to do more.