AINewsnow

I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

There are good benchmarks for role-play models. RoleLLM, PingPong, LoCoMo, LongMemEval. All of them test raw models. None of them test the apps people actually download. That gap matters more than it sounds. Replika, Character.AI, Nomi, Kindroid, Talkie, and the AI dating simulators all ship a mode…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-03 10:27 · DEV Community — AI
    I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam

More stories

  1. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Introducing Astra for Law — OpenAI News
  4. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  5. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  6. What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog
  7. Novo Nordisk Will Use Anthropic’s Claude for Drug Research — Wall Street Journal Technology
  8. Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →