The agent is average on purpose.
Qwack is the agent on every Ducker challenge, and it is deliberately not a flagship. It writes code that usually runs. It is lazy, it will not go looking for your edge cases, it sometimes over-builds a simple change, and it can be confidently wrong about the codebase it is working in. Ask it properly and it does excellent work.
Same model, same behaviour, for every candidate you invite.
Strong engineers still ship strong code with it
Qwack does not know which failure mode matters here, and it will not raise one on its own. Candidates who read what came back, worked out the side effects, asked for the missing test and pushed until they would defend the result submit code that passes review. Candidates who keep the first thing that runs submit code that does not.
- It writes what it was asked for and stops — the missing transaction stays missing until someone notices
- Told to review, harden, refactor, secure or test the implementation, it does high-quality work
- It can be wrong about the repository it is editing, so the candidate has to actually read the diff
Qwack has no opinion about your conventions
It will reach for a string union where an enum belongs, leave a module-level array next to a class, name a parameter d, and hand the caller an any. All of it runs. None of it is what you want in the repository six months from now. Candidates who care about design patterns, naming and encapsulation fix it before they submit — and that is what the code review grades.
12- const STATUS_OPTIONS = ['connected', 'disconnected'];
13- type S = 'connected' | 'disconnected';
14- export function proc(d: any, s: S) {
15- if (s === 'connected') return d.map((x: any) => x.v);
16- }
12+ export enum ConnectionStatus {
13+ Connected = 'connected',
14+ Disconnected = 'disconnected',
15+ }
17+ export class ReadingCollector {
18+ private readonly liveStatus = ConnectionStatus.Connected;
20+ collect(samples: Sample[]): number[] {
21+ return this.isLive()
22+ ? samples.map((sample) => sample.value)
23+ : [];
24+ }
25 }How Qwack behaves
It implements, it does not improve
Ask for the endpoint and you get a working endpoint. No validation nobody requested, no failure path volunteered, no test for the case the brief did not mention.
Ask properly and it delivers
Told to handle concurrency, cover the failure path or clean up the data model, Qwack does the work well. The ceiling is set by the candidate, not by the model.
It gets things wrong
It can misread the codebase, over-engineer a two-line change, or ignore a convention the repository already follows. Accepting its output unread is how a submission fails.
Why not hand everyone a frontier model
A model that quietly fixes everything hands every candidate the same submission. Qwack stops at "this runs", so the distance between two engineers shows up in the code they turn in.
- 1
Everyone gets the same collaborator
Same model, same laziness, same blind spots. Differences in the submission come from the person, not from whose AI subscription is better.
- 2
The gap is the exercise
Code that runs but is not code you would own is exactly what engineers get from agents every day. Closing that gap is the task.
- 3
The grade comes from the code
Hidden tests and an agentic code review judge the submission. How many prompts it took, and who typed which line, do not enter the score.
See what a good engineer gets out of an average agent.
Send one challenge and read the code that comes back.
Get early accessNo credit card · Pay only for assessments candidates start