I'm not going to write some big blog post here, but this chat log is quite entertaining. I wrote a fundamental string algorithm in C and GPT 5.6 Sol on medium effort hallucinated not one but two correctness issues with the code!
Maybe this is not surprising to others but it's been a long while since I caught one of these models making such a glaring error.
And maybe it's just the harness and the backing model is mostly the same. But I would expect Codex to give better responses to code questions than chatGpt using the "same" model.
I went back through Azure foundry, and it looks like they have gotten rid of reasoning variants, so might just be one model now.
I'm not going to write some big blog post here, but this chat log is quite entertaining. I wrote a fundamental string algorithm in C and GPT 5.6 Sol on medium effort hallucinated not one but two correctness issues with the code!
Maybe this is not surprising to others but it's been a long while since I caught one of these models making such a glaring error.
I'm pretty sure there is a completely different model for Chat vs Coding.
Like if you deploy 5.6 on azure foundry there is a chat model and a coding/reasoning model
I learned something new today
And maybe it's just the harness and the backing model is mostly the same. But I would expect Codex to give better responses to code questions than chatGpt using the "same" model.
I went back through Azure foundry, and it looks like they have gotten rid of reasoning variants, so might just be one model now.