FILTERED RESULTS
FILTERS
Ads Top
DARK MODE
CHART
MCap $2.7T +0.1%24h Vol $53.8B -51%Fear & Greed 63/100Alts Index 35/100
BTC.D 58.3% -0.1%Stable.D 10.0% 0%ETH.D 11.6% +0.1%Others.D 20.1% 0%
AI$0.3437+36.42%UAI$0.7955+18.86%PROM$5.778+14.88%BTW$0.5531+10.54%ZRX$0.1095+9.35%ETHFI$0.7402+8.52%PENDLE$2.185+8.26%PUMP$0.00387349+8.17%WLFI$0.0575+7.96%PYTH$0.0552+7.44%
APEPE$0.00000133-20.5%MET$0.2317-13.26%Q$0.0239-11.18%HNT$0.4763-8.96%NEAR$2.364-7.25%MARSCOIN$0.1171-4.92%MORPHO$2.229-4.33%RAIN$0.0152-4.23%LIT$4.320-3.98%FF$0.1495-3.67%
Top movers 24h
    Filters
      Coins
      Sentiment
      Impact
      Search
      FILTERED RESULTS

        

      Upgrade your plan
      Dashboard

      Microsoft Launches ThinkingBox to Evaluate AI Agent Reliability

      Microsoft has unveiled ThinkingBox, an open-source framework aimed at assessing the reliability of AI agents in performing business tasks. Announced on August 19, 2026, by Principal Machine Learning Engineer Liang-Chun Tsai, the framework offers a new approach by focusing on actual changes made in a database rather than relying on traditional transcript analysis.

      ThinkingBox is accompanied by a benchmark called ThinkingBox-Bench, which revealed concerning results during initial tests. Microsoft evaluated 12 different AI models across 507 tasks in five business domains. The best-performing model achieved a 65.36% success rate on its first attempt, but only 25.25% across 20 trials, highlighting what Microsoft terms the 'discovery-reliability gap.' This gap indicates that even the most effective AI agents struggle to deliver consistent results over repeated attempts.

      The framework operates by creating isolated environments for each trial, allowing for a rigorous evaluation of the AI's performance. This method exposes failures that traditional evaluation techniques might overlook, such as instances where an agent appears to complete a task without actually achieving the desired outcome. Microsoft aims for ThinkingBox to set a new industry standard for evaluating AI reliability, making the framework and its benchmark available on GitHub.

      © 2026 KLEA News. All Rights Reserved. This article is provided for informational purposes only. It is not offered or intended to be used as legal, tax, investment, financial, or other advice.

      Source: KLEA News

      .

      Terra Founder Do Kwon Sentenced to 15 Years in Prison for Fraud