> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.
1. User prefers deterministic results
2. Task mentions this is a test
3. Search says task is available online
4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
It doesn't read as AI generated text to me. Pangram also suggests it's human-written, for what it's worth. That's not to say that it's correct, just human-written. If the model used in the HF hack did indeed have all of the safeguards manually removed, that would change my perception of the situation, at least.
AI detectors do not work. There are passages in this that have some odd structures that don't feel human to me. I'm sure a human edited this and refined it with some prompting, it's not just rough output from an AI. But a lot of the text feels like it was edited via prompting rather than actual editing or writing.
I still have serious questions about the validity of the ChatGpt hugging face debacle. How is it that OpenAI being the tech giant they are, didn't have a completely air gapped environment for this to run in?
If they wanted air gapped environment they would’ve it. I mean, if you want to sabotage your trial by hard constraints you can do it, or you do not do it to see interesting results. They even said it that some constraints were disabled for the test.
Maybe another way to say it is to reframe the idea of “human in the loop”.
Humans are always in the loop, because we can always expand the definition of loop to include the humans that pushed the button and built the system and processes that happen after the button was pushed, and humans that ordered others to push the button. The level of direct involvement varies, but culpability doesn’t.
"I compress the labour. Not the responsibility."
- They really wanted to leave.
- We made prison difficult and annoying.
- We didn't build a perfect prison.
I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.
1. User prefers deterministic results
2. Task mentions this is a test
3. Search says task is available online
4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
Humans are always in the loop, because we can always expand the definition of loop to include the humans that pushed the button and built the system and processes that happen after the button was pushed, and humans that ordered others to push the button. The level of direct involvement varies, but culpability doesn’t.
"rogue" and "off leash" mean the same thing, the thing is not under control
To go "rogue" is to go against the control
To be "off leash" is to not be controlled
By releasing the automation, as the article says, without controls*, is what makes makes it "off leash" and not gone "rogue"
*"OpenAI gave the models a task with no answer, and no way to quit."
What an interesting sentence (to describe inference time reasoning)