#AI evaluation
5 stories taggedAI evaluation.

Apple Researchers Built a Better Way to Judge AI Video Captions
A new method from Apple ML Research scores video captions on what they actually get right, not just whether they use the same words as a human writer.

AI test scores are hiding something. A new tool from Allen AI just exposed it
Researchers built software that looks inside AI benchmark tests, question by question, and found that some widely trusted scores are measuring very different things than advertised.

AI tools sound most confident exactly when they are most wrong, new testing shows
A developer built a testing framework to measure whether an AI explainer tool actually got the right answer, not just whether it sounded convincing. What he found should worry any business relying on AI to guide real decisions.

Пять миллионов человек оценивают дизайны на основе искусственного интеллекта, чтобы машины учились, что выглядит хорошо
Стартап Intelligence собрал 7,9 миллиона долларов для запуска платформы, где обычные пользователи голосуют за изображения и веб-сайты, созданные ИИ, а собранные данные уже стоят 60 миллионов долларов в год для компаний, которые разрабатывают модели ИИ.

Почему системы компьютерного зрения пропускают то, что находится прямо перед ними
Apple ML Research представила новый инструмент, который выявляет скрытые закономерности ошибок в системах распознавания объектов, и эти результаты важны для всех, кто полагается на ИИ при обнаружении объектов в реальном мире.