The problem was silent
Since late September, agents connected to our platform were starting work but never finishing. Teams would see runs stuck in progress, no error message, no visibility into what went wrong. The agent had called the model, pulled in the tools, then hung.
The real failure was buried in server logs: the checkpoint serializer couldn't store the agent's intermediate state. Every time the agent hit a certain execution pattern, the state format didn't match what the platform's validator expected. The worker process would exit, the run would stall, and the team got silence.
What we shipped
We bumped the agent runtime to LangGraph 1.4.12 and updated the SDK dependencies to match. That version uses the correct checkpoint format. The platform validator now accepts the state, the worker stays alive, and agents finish.
We ran 27 agent tests. All pass. We ran 704 dashboard unit tests. All pass. We tested the exact serialization format against the production server validator. Checkpoints that failed before now validate clean.
What changed for your team
If you're connected to our chat agent, widget runner, or commitment extractor, runs now complete end-to-end. No more hanging. No more silent failures. The agent does the work and returns a result. 💧