MEDIUMAi
Global
Sol Searching | Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?
·Source: SentinelLabs
Updated:
Executive Summary
A real-world benchmark tests whether powerful AI models can keep an investigation trustworthy when new evidence invalidates their conclusions.
Analysis
A real-world benchmark tests whether powerful AI models can keep an investigation trustworthy when new evidence invalidates their conclusions.