
Technik
29.5.2026KI-Zusammenfassung
AI Personal Assistants Struggle with Real-World Complexity in New Benchmark
A new benchmark, Claw-Anything, tests AI personal assistants on realistic, long-horizon tasks across multiple devices and services, revealing poor performance (e.g., GPT-5.5 scored 34.5%) due to complexity and noise in simulated user activity.
D
Decrypt4 Min. Lesezeit