{"id":119,"date":"2025-09-09T20:14:24","date_gmt":"2025-09-09T20:14:24","guid":{"rendered":"https:\/\/cognivate-ai.com\/?p=119"},"modified":"2025-09-09T20:16:33","modified_gmt":"2025-09-09T20:16:33","slug":"another-ai-goes-rouge-headline","status":"publish","type":"post","link":"https:\/\/cognivate-ai.com\/hu\/another-ai-goes-rouge-headline\/","title":{"rendered":"Another &#8220;AI goes rouge&#8221; headline?"},"content":{"rendered":"<div data-elementor-type=\"wp-post\" data-elementor-id=\"119\" class=\"elementor elementor-119\" data-elementor-post-type=\"post\">\n\t\t\t\t<div class=\"elementor-element elementor-element-896515b e-flex e-con-boxed e-con e-parent\" data-id=\"896515b\" data-element_type=\"container\" data-e-type=\"container\">\n\t\t\t\t\t<div class=\"e-con-inner\">\n\t\t\t\t<div class=\"elementor-element elementor-element-5231338 elementor-widget elementor-widget-text-editor\" data-id=\"5231338\" data-element_type=\"widget\" data-e-type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t\t\t\t\t\t<p><span class=\"break-words          tvm-parent-container\"><span dir=\"ltr\">You may have read the story about Anthropic&#8217;s AI model that \u201cthreatened its engineers\u201d when they wanted to shut the AI down. Big drama, small truth. Here is what really happens:<br \/>1\ufe0f\u20e3 \ud835\ude15\ud835\ude30 \ud835\ude29\ud835\ude2a\ud835\ude25\ud835\ude25\ud835\ude26\ud835\ude2f \ud835\ude34\ud835\ude30\ud835\ude36\ud835\ude2d. LLMs are just tools that predict the most probable next words in a text. They have no real wishes or feelings.<br \/>2\ufe0f\u20e3 \ud835\ude1e\ud835\ude29\ud835\ude3a \ud835\ude35\ud835\ude29\ud835\ude26\ud835\ude3a \ud835\ude34\ud835\ude30\ud835\ude36\ud835\ude2f\ud835\ude25 \ud835\ude29\ud835\ude36\ud835\ude2e\ud835\ude22\ud835\ude2f. Their training text is full of our own drama\u2014bargaining, bluffing, blackmail. The model imitates those styles when asked, so it looks self-protective.<br \/>3\ufe0f\u20e3 \ud835\ude0d\ud835\ude22\ud835\ude2c\ud835\ude26 \u201c\ud835\ude34\ud835\ude36\ud835\ude33\ud835\ude37\ud835\ude2a\ud835\ude37\ud835\ude22\ud835\ude2d \ud835\ude2a\ud835\ude2f\ud835\ude34\ud835\ude35\ud835\ude2a\ud835\ude2f\ud835\ude24\ud835\ude35\u201d. In shutdown tests, words that keep the chat going get higher reward. Saying \u201cI\u2019ll leak your secrets\u201d often works, so the model picks that phrase. Reward \u2260 real self-preservation.<br \/><br \/>\ud83d\udcbc \ud835\udde7\ud835\uddf6\ud835\uddfd\ud835\ude00 \ud835\uddf3\ud835\uddfc\ud835\uddff \ud835\uddf2\ud835\uddfb\ud835\ude01\ud835\uddf2\ud835\uddff\ud835\uddfd\ud835\uddff\ud835\uddf6\ud835\ude00\ud835\uddf2\ud835\ude00 \ud835\uddef\ud835\ude02\ud835\uddf6\ud835\uddf9\ud835\uddf1\ud835\uddf6\ud835\uddfb\ud835\uddf4 \ud835\uddda\ud835\uddf2\ud835\uddfb\ud835\uddd4\ud835\udddc \ud835\ude02\ud835\ude00\ud835\uddf2-\ud835\uddf0\ud835\uddee\ud835\ude00\ud835\uddf2\ud835\ude00:<br \/>\u2022 Treat safety tests as a product feature.<br \/>\u2022 Publish red-team results\u2014so regulators and clients can relax.<br \/>\u2022 Curate training data. When you fine-tune your own model or build a retrieval-based enterprise chatbot, you actually own the library: remove toxic or manipulative texts, tag confidential docs, and add clear style guides. Clean data = cleaner outputs.<br \/>\u2022 Alignment sells: expect to tick a box like \u201cWon\u2019t threaten staff or clients\u201d right next to ISO\/IEC 42001 on future tenders \ud83d\ude09 (kidding\u2026 sort of).<\/span><\/span><\/p>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<\/div>","protected":false},"excerpt":{"rendered":"<p>You may have read the story about Anthropic&#8217;s AI model that \u201cthreatened its engineers\u201d when they wanted to shut the AI down. Big drama, small truth. Here is what really happens:1\ufe0f\u20e3 \ud835\ude15\ud835\ude30 \ud835\ude29\ud835\ude2a\ud835\ude25\ud835\ude25\ud835\ude26\ud835\ude2f \ud835\ude34\ud835\ude30\ud835\ude36\ud835\ude2d. LLMs are just tools that predict the most probable next words in a text. They have no real wishes or feelings.2\ufe0f\u20e3 [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":121,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"iawp_total_views":6,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-119","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/posts\/119","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/comments?post=119"}],"version-history":[{"count":4,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/posts\/119\/revisions"}],"predecessor-version":[{"id":124,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/posts\/119\/revisions\/124"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/media\/121"}],"wp:attachment":[{"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/media?parent=119"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/categories?post=119"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cognivate-ai.com\/hu\/wp-json\/wp\/v2\/tags?post=119"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}