I’ve been testing GPT-6 Sol in ChatGPT Work across real development projects, and I keep seeing more execution mistakes than I do with Astra.
Missed instructions. Incomplete gates. Context that gets read but does not always survive the full execution.
I went looking to see if it was just me. It isn’t. Other users are reporting similar problems with context handling and long-running Work/Codex sessions.
That does not mean Sol is universally worse. It does mean I’m changing how I use it.
For technical project work, I’m back to Astra first, usually at the lowest reasoning level that can reliably finish the job.
Sol still has a place.
It just is not my default for technical Work anymore.
Loading comments...