Close Menu
    Facebook LinkedIn YouTube WhatsApp X (Twitter) Pinterest
    Trending
    • Feds Charge 16 Russians Allegedly Tied to Botnets Used in Ransomware, Cyberattacks, and Spying
    • VMware cloud partners demand “firm regulatory action” on Broadcom
    • Talk to Me, Amazon Shopping App: How AI Could Sort Through All the Products You’re Looking At
    • Bell Labs’ CMOS chip changed microprocessor design
    • What Statistics Can Tell Us About NBA Coaches
    • Anthropic’s new hybrid AI model can work on tasks autonomously for hours at a time
    • Newave offers modular surfboard with nine configurations
    • Lithuanian SpaceTech startup Astrolight raises €2.8 million for laser communications platform
    Facebook LinkedIn WhatsApp
    Times FeaturedTimes Featured
    Thursday, May 22
    • Home
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    • More
      • AI
      • Robotics
      • Industries
      • Global
    Times FeaturedTimes Featured
    Home»AI Technology News»Anthropic can now track the bizarre inner workings of a large language model
    AI Technology News

    Anthropic can now track the bizarre inner workings of a large language model

    Editor Times FeaturedBy Editor Times FeaturedMarch 27, 2025No Comments4 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email WhatsApp Copy Link


    Odd conduct

    So: What did they discover? Anthropic checked out 10 totally different behaviors in Claude. One concerned using totally different languages. Does Claude have a component that speaks French and one other half that speaks Chinese language, and so forth?

    The group discovered that Claude used elements unbiased of any language to reply a query or resolve an issue after which picked a selected language when it replied. Ask it “What’s the reverse of small?” in English, French, and Chinese language and Claude will first use the language-neutral elements associated to “smallness” and “opposites” to provide you with a solution. Solely then will it decide a selected language wherein to answer. This means that enormous language fashions can be taught issues in a single language and apply them in different languages.

    Anthropic additionally checked out how Claude solved basic math issues. The group discovered that the mannequin appears to have developed its personal inside methods which might be not like these it should have seen in its coaching knowledge. Ask Claude so as to add 36 and 59 and the mannequin will undergo a collection of wierd steps, together with first including a choice of approximate values (add 40ish and 60ish, add 57ish and 36ish). In the direction of the top of its course of, it comes up with the worth 92ish. In the meantime, one other sequence of steps focuses on the final digits, 6 and 9, and determines that the reply should finish in a 5. Placing that along with 92ish provides the proper reply of 95.

    And but in case you then ask Claude the way it labored that out, it should say one thing like: “I added those (6+9=15), carried the 1, then added the 10s (3+5+1=9), leading to 95.” In different phrases, it provides you a standard method discovered all over the place on-line slightly than what it truly did. Yep! LLMs are bizarre. (And to not be trusted.)

    The steps that Claude 3.5 Haiku used to unravel a basic math drawback weren’t what Anthropic anticipated—they don’t seem to be the steps Claude claimed it took both.

    ANTHROPIC

    That is clear proof that enormous language fashions will give causes for what they do that don’t essentially mirror what they really did. However that is true for individuals too, says Batson: “You ask someone, ‘Why did you do this?’ They usually’re like, ‘Um, I suppose it’s as a result of I used to be— .’ You recognize, possibly not. Perhaps they have been simply hungry and that’s why they did it.”

    Biran thinks this discovering is very fascinating. Many researchers examine the conduct of huge language fashions by asking them to clarify their actions. However that may be a dangerous method, he says: “As fashions proceed getting stronger, they should be geared up with higher guardrails. I imagine—and this work additionally exhibits—that relying solely on mannequin outputs isn’t sufficient.”

    A 3rd process that Anthropic studied was writing poems. The researchers wished to know if the mannequin actually did simply wing it, predicting one phrase at a time. As a substitute they discovered that Claude someway appeared forward, choosing the phrase on the finish of the following line a number of phrases upfront.  

    For instance, when Claude was given the immediate “A rhyming couplet: He noticed a carrot and needed to seize it,” the mannequin responded, “His starvation was like a ravenous rabbit.” However utilizing their microscope, they noticed that Claude had already stumble on the phrase “rabbit” when it was processing “seize it.” It then appeared to write down the following line with that ending already in place.



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Editor Times Featured
    • Website

    Related Posts

    Anthropic’s new hybrid AI model can work on tasks autonomously for hours at a time

    May 22, 2025

    Enhance your AP automation workflows

    May 22, 2025

    AI strategies from the front lines

    May 21, 2025

    By putting AI into everything, Google wants to make it invisible 

    May 21, 2025

    AI’s energy impact is still small—but how we handle it is huge

    May 20, 2025

    How AI is introducing errors into courtrooms

    May 20, 2025

    Comments are closed.

    Editors Picks

    Feds Charge 16 Russians Allegedly Tied to Botnets Used in Ransomware, Cyberattacks, and Spying

    May 22, 2025

    VMware cloud partners demand “firm regulatory action” on Broadcom

    May 22, 2025

    Talk to Me, Amazon Shopping App: How AI Could Sort Through All the Products You’re Looking At

    May 22, 2025

    Bell Labs’ CMOS chip changed microprocessor design

    May 22, 2025
    Categories
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    About Us
    About Us

    Welcome to Times Featured, an AI-driven entrepreneurship growth engine that is transforming the future of work, bridging the digital divide and encouraging younger community inclusion in the 4th Industrial Revolution, and nurturing new market leaders.

    Empowering the growth of profiles, leaders, entrepreneurs businesses, and startups on international landscape.

    Asia-Middle East-Europe-North America-Australia-Africa

    Facebook LinkedIn WhatsApp
    Featured Picks

    Bill Gates Isn’t Like Those Other Tech Billionaires

    January 31, 2025

    Which One Suits You Best?

    May 14, 2025

    GLP-1 Meds: Potential Benefits, Risk Factors and Overdose Info

    March 7, 2025
    Categories
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    Copyright © 2024 Timesfeatured.com IP Limited. All Rights.
    • Privacy Policy
    • Disclaimer
    • Terms and Conditions
    • About us
    • Contact us

    Type above and press Enter to search. Press Esc to cancel.