Prompted by a real incident: an MSBuild node-reuse worker (its own
/nodeReuse:true default) survived a build getting interrupted, came
back deadlocked, and got reused by the next `msbuild` invocation -
which then failed with a confusing, unrelated-looking
`System.TypeLoadException` on Microsoft.VisualStudio.Telemetry on
every call for hours, until the stale process was killed by hand.
That's exactly the "unrelated blocker" noted in this repo's own
earlier session notes (CLAUDE.md) while debugging a real project's
build - it wasn't a missing dependency, it was a corrupted reused
process.
Two changes:
- vintner now forces /nodeReuse:false on every msbuild invocation
(unless the caller already passed their own /nodeReuse or /nr
switch), so a wedged worker can never poison a later, unrelated
build in the first place. Costs each invocation the couple-hundred-
ms/node startup time node reuse exists to save.
- VINTNER_TIMEOUT (a duration string, e.g. "30m") bounds how long any
single tool invocation is allowed to run, for the case something
wedges that isn't MSBuild-specific. Every exec.Command site in
internal/wrapper now goes through a shared newToolCommand
constructor that, when the timeout is set, kills the *whole*
process group (not just the immediate `wine` process - a wedged
child surviving under it is exactly the scenario this needs to
reach) via a context deadline, and reports a clear "timed out after
Xm" message (exit 124, matching the timeout(1) convention) instead
of a bare "signal: killed". Unset by default - every real build
observed stays unbounded, matching Windows' own behavior.
Verified end-to-end, not just at the unit level: a real `sleep 30`
through the `cmd` native wrapper with VINTNER_TIMEOUT=1s was killed
within the deadline and reported the timeout clearly (exit 124); a
real `cl` invocation with the same 1s timeout finished normally
(0.26s) without being mistaken for a hang.