• jj4211@lemmy.world
    link
    fedilink
    English
    arrow-up
    8
    ·
    14 hours ago

    Generally speaking, you can get an AI to admit fault and accept blame for something that never even happened.

    It reacts to the operator input that something happened that warrants blame. So the natural continuation of that conversation is to acknowledge the situation, accept blame, and promise to do better. It generally doesn’t model what did or did not happen, nor does the promise to do better mean anything. It does weigh in the context window that it wouldn’t be a natural continuation to do that action after text saying that it won’t happen, but can evaporate and doesn’t carry over when the operator goes to a new context window, which of course is a crazy nuance for the rando user to understand.

    • Lovable Sidekick@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 hours ago

      Hopefully most users get that a promise AI makes won’t stay in force forever, but it should definitely last through the current conversation. I mean, I would call that a minimal performance standard, and if it’s not true the AI shouldn’t make promises in the first place.

      I’m a little hazy on the part where accepting blame for something that didn’t happen is a natrual continuation of a conversation. It’s certainly not natural in human conversations. I would expect a well written and well trained AI to correct false assumptions or factual errors made by the user, not pander to them - unless it’s been told never to contradict the user.