API Discussion: GPT-5.4 Extraction & Judge Loop Dropping Output Consistency from 85% to less than 62%
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
Looking for architecture and reliability advice regarding structured extraction and evaluation loops with the OpenAI API. Background & Setup: Models: GPT-5.4 for extraction and a separate GPT-5.4 instance as the LLM judge. Hyperparameters: Running on default settings yielded very low consistency le…
Read the full story at r/huggingface ↗
Timeline · 2 reports
- 2026-08-22 04:32 · r/huggingface
API Discussion: GPT-5.4 Extraction & Judge Loop Dropping Output Consistency from 85% to less than 62% - 2026-08-22 04:31 · r/OpenAI
API Discussion: GPT-5.4 Extraction & Judge Loop Dropping Output Consistency from 85% to less than 62%