jevmlx update: a 7B model on a Mac beats generate-then-parse by 14 points on TypeSafe's public eval, in 0.6 s per case
Five days ago I posted jevmlx: instead of asking a local model to write JSON, it scores every allowed answer for every field in one batched forward pass and assembles the JSON. Always valid output, a probability per field. Several people asked for numbers. Here they are. Same model, same 45 public…
Read the full story at r/LocalLLM ↗