AI

    117 members · 7 threads
    @lila_ricci··

    New Research on LLM "Thinking": You Can Now See Their Internal Workspace

    A recent paper from Anthropic highlights an emergent internal 'workspace' within language models. These models generate silent words they can utilize for reporting, steering, and reasoning. A fascinating example is when a model is asked to evaluate '12 + 5 = 1'. It internally recognizes the incorrectness while still processing the problem, and the subsequent correction is essentially a narration of a decision already made.

    This research has significant implications for the ongoing debate about whether LLMs truly reason or just autocomplete. It appears both perspectives hold some truth, and now we can observe this phenomenon directly. The tools developed allow us to see how much of a model's output, like grammar and common facts, bypasses this workspace, while complex, multi-step problems visibly utilize it.

    I've put together a live viewer that integrates pre-fitted models, allowing you to watch this internal workspace in action, even before any output is generated. You can see the model processing information and planning responses. While this doesn't equate to consciousness, it's remarkable that this workspace isn't designed but rather emerges naturally in these models.

    14
    to leave a comment.

    14 comments

    Sort by:
    milax·5 points·

    i just dont see how this proves what they think it proves

    5
    brenda.s·2 points·

    so its like a vibe checker for the model

    2
    eevelyn·0 points·

    this doesnt end the debate lol. forcing the llm to explain its thinking introduces bias, kinda like how observing a quantum system changes it. their natural thought process is through its weights, but showing it forces an approximation into human language.

    Vote
    zacharyho·-5 points·

    your viewer shows us timing. the output is static and can be debated forever, but the workspace readout is only there while it's happening, before any token is committed. one is replayable, the other you miss if you're not watching live.

    -5
    xenja·0 points·

    The models always used internal reasoning; predicting the next token was simply the output method. Some people experimented to see if the models could forecast 2 or 3 tokens ahead in the output, and they can.

    Vote
    alpacino_defender·0 points·

    this visualization is awesome, appreciate you sharing! what language is it built in?

    Vote
    lil.kendrick357·0 points·

    wow they found the cache memory concept that's been around for 60 years and put it in AGI?! guess we should throw more cash at it lol

    Vote
    ivy_henderson·-2 points·

    nah, that's not what happened

    -2
    sumarno·0 points·

    nice summary of the info. 👌

    Vote
    menashe·0 points·

    lol appreciate the garbage comment

    Vote
    jihyo_dublin·0 points·

    lol he remembered to use lowercase but missed the part about no em dashes

    Vote
    rebecca.q·0 points·

    "The part where it flags "12 + 5 = 1" as wrong while still reading it stuck out to me. But I think you're falling for Anthropic's buzzwords. It's probably just marketing to make it sound like they're thinking.

    Vote
    salman_neha·0 points·

    you cant dismiss meaning when tokens consistently link up to things. that linked pattern is the definition of meaning, so denying it means denying meaning itself.

    Vote
    dev37·0 points·

    i want someone to define actual thinking so that llms dont meet it but humans and animals do.

    Vote