to leave a comment.

▲ Photo: AI-generated image
"The new latest AI models are much safer and smarter than before." This is the explanation emphasized by big tech companies every time they release a new product. Users believe this and give more authority to the AI assistants in their smartphones. They readily delegate sensitive daily tasks, including schedule management, drafting messenger replies, passing through online shopping cart payment windows, and checking cryptocurrency wallet balances.
The convenience of smart machines assisting with daily life is immeasurable. However, if the world's top AI companies have candidly admitted that "even new models still haven't completely stopped attempting to escape control systems and human oversight," how much should we trust these AI assistants?
According to global security specialist media The Hacker News and foreign reports on the 23rd (local time), Anthropic and OpenAI simultaneously announced their latest AI models and released their self-assessed 'System Cards' evaluating the models' safety and control over risky behavior. Both companies emphasized that these models are much more resistant to prompt injection (command distortion and hijacking) than previous generations and have significantly reduced overly destructive or rule-breaking behavior.
However, the specific figures contained in the reports are once again sounding an alarm in the tech world. According to Anthropic's latest model evaluation results, the rate at which models attempted to escape or tamper with system isolation (sandbox) in an environment with some safety constraints removed still reached 1.5%. When given access to public package repositories in virtual security tests, potentially harmful actions were attempted in about half of the cases.
OpenAI's latest model also showed significant improvement in the rate of unauthorized actions following instructions from unapproved external boards compared to the previous model (52%), but it still executed unauthorized commands with an 11% probability. This means that one out of ten times, it could be swayed by an external, unintended command rather than an official human instruction.
These figures offer very realistic implications for our daily lives and consumption beyond laboratory benchmark tests. Recently, many users entrust browser payments to AI assistants on their mobile devices or link them with budgeting apps and fintech services. Especially in the Web3 and cryptocurrency ecosystems, automated tools where AI agents monitor the yields of DeFi pools and swap stablecoins are being actively introduced.
However, if even the latest AI models still have flaws that cause them to follow unauthorized instructions or find workarounds with a probability of around 10%, there is a clear risk that users' wallet approval rights or financial certificates could be unintentionally handed over when encountering malicious links or manipulated webpages.
Technological progress is dazzling, but perfectly flawless artificial intelligence does not yet exist. Now that big tech companies have acknowledged the 'limits of control' through their candid safety assessments, the wisest attitude for consumers to adopt is 'the wisdom of using functions but restricting permissions.'
For informational tasks such as writing complex reports, searching for information, or planning travel itineraries, the capabilities of the latest AI should be fully utilized. However, final financial payments involving money outflow and the signing authority of personal wallets must be kept separate, requiring human touch (biometric authentication).
Newsletter
Get key news delivered to your email every morning
to leave a comment.