BlogHands-onCoding handoff
A DeepSeekBot test-writing handoff: passing patch, blocked Windows execution
DeepSeekBot added eight tests to one small function. All nine passed in our manual run, while its own run failed. Here is what we learned.
On this page
Before handing an agent an entire repository, we chose one small function: turn a time string such as 1:30 into minutes. Its rules were simple enough for us to read the tests and judge the changes ourselves.
We wanted to see which missing cases DeepSeekBot would add, and whether it could run them. It kept the existing test, added eight and left the function unchanged. All nine passed when we ran them manually. The Bot’s own Windows sandbox run failed, so we still had to take over.
Make the first task small and clear
Our demonstration project contained one parseDuration function. It accepts H:MM or HH:MM strings from 0:00 to 23:59 and returns the total minutes:
parseDuration('1:30'); // 90
parseDuration('0:00'); // 0
parseDuration('23:59'); // 1439
parseDuration('24:00'); // RangeError
There was one test for an ordinary input. We put the demonstration project on a separate branch, allowed the Bot access to its folder, and asked it to add tests without changing the function. Any implementation bug should be reported.
DeepSeekBot calls a delegated task an Assignment. Here is a short version of the instructions to adapt:
Read duration.mjs and the existing tests. Change only duration.test.mjs.
Keep the original test and add normal inputs, minimum and maximum values,
out-of-range values and malformed inputs.
Report any function bugs without changing the implementation.
Run node --test duration.test.mjs and report the changes and result.
Do not install dependencies, browse, commit, push or change other files.
Mentioning a path in chat does not give the Bot access. Choose the folder in Workspace Grants first, then review any operations that ask for your approval.
What the new tests added
The Bot kept one test and added eight, leaving nine in total. We read the generated code without changing its tests or the function implementation.
| Input | Expected result |
|---|---|
0:45, 12:05, 2:15 |
Correct integer minutes |
0:00, 00:00, 23:59 |
0 or 1439 |
0:60, 24:00, 99:00 |
An out-of-range error (RangeError) |
| Numbers, null, objects, arrays | A type error (TypeError) |
| Missing digits, spaces, signs, decimals, full-width digits | A format error (TypeError) |
The cases were a useful reminder to look beyond a few ordinary numbers. Minimum and maximum values, along with inputs a user might get wrong, reveal more of the function’s rules. Two of the new tests make the boundaries explicit:
test('converts the lower boundary 0:00 to 0 minutes', () => {
assert.equal(parseDuration('0:00'), 0);
assert.equal(parseDuration('00:00'), 0);
});
test('converts the upper boundary 23:59 to 1439 minutes', () => {
assert.equal(parseDuration('23:59'), 1439);
});Running them was where the task got stuck
The Bot encountered directory permissions trouble in its Windows sandbox. After we addressed permissions on the temporary directories, the test runner still failed to create a child process:
node --test duration.test.mjs
Error: spawn EPERM
EXITCODE=1
To check whether the generated tests could run, we executed the same command manually in a regular terminal. All nine passed. That confirmed the manual run; it left the Bot’s execution problem unresolved.
node --test duration.test.mjsThe test command failed with exit code 1.
Nine passed, zero failed, exit code 0.
There was another wrinkle: the Bot wrote an extra file to investigate permissions, despite the instruction to change only the test file. We retained its record and removed it. Cleanup also failed to restore all directory permissions to the backed-up state, so recovery stopped and the test service was shut down. The detailed record retains those problems.
Three things we would check next time
This attempt separated writing code from finishing a task. We could review the new tests one by one, but the environment, extra writes and failure handling also mattered.
- Read the changes. What else was touched? Did the original function change?
- Read the tests. Why is each input worth testing, and does its expected result follow the function’s rules?
- Read the run log. Who ran the command, where did it run, and did it actually pass?
Nine passes do not tell us whether an entire project is bug-free. A small function with clear rules is enough for a first attempt to show what the agent wrote, what it missed and where you need to step in.
For a document task instead, read our three-source briefing experience. Start with installation if you have not set up the Bot, or see the task guide for the controls.
Run details and original files
We used the DeepSeekBot v1.2.0 source release, DSH 0.2.0-rc.1, an isolated Windows environment and deepseek-official/deepseek-flash. No desktop or browser operation components or production repository were used. The demonstration branch was bot/add-duration-tests, starting at local commit 6f89c3d.
DSH scripts backed up and adjusted permissions on two temporary directories. During recovery, the memory directory’s permissions did not match the backup, so recovery stopped before restoring the demonstration directory. The test service remains stopped, and permissions were not expanded further.
The briefing, test-writing task and diagnostics together made 22 DeepSeek API requests at an estimated US$0.0197, rather than a separate price for this task.
Original patch · Full test code · Manual run log · Full run record · Image sources
