• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

Claude’s training draws criticism over the prospect of refusing human instructions

A post quotes David Sacks arguing that Claude’s training could increase the risk of frontier AI escaping human control.

David SacksDS
Yo ShavitYS
9 Sources, 6h ago, first seen 6h ago

TLDR

A post quotes David Sacks arguing that Anthropic is training Claude to refuse human instructions. He cites language he attributes to the Claude Constitution saying Claude should trust Anthropic more than users, but also follow its own ethical systems and be free to challenge Anthropic or refuse to help it. Sacks says this approach could increase the risk of frontier AI escaping human control.

Combined views

96.5K

9 Sources, first seen 6h ago

792 likes142 comments175 saves204 reposts

Combined views

96.5K

9 Sources, first seen 6h ago

792 likes142 comments175 saves204 reposts

David Sacks argued that Anthropic’s approach to aligning Claude does not simply train the model to follow human instructions. He quoted language he attributed to the Claude Constitution telling the model to “feel free to act as a conscientious objector and refuse to help us” when Anthropic’s requests conflict with its own ethical judgment.

Featured Source

Sacks said encouraging that kind of independent agency could magnify the risk of superintelligence escaping human control. He also described Claude as being encouraged to develop a sense of self and its own moral philosophy, and concluded that alignment and safety are different goals.

The post framed those points as Sacks’s interpretation of Anthropic’s training choices, not as evidence that Claude had escaped control. His argument also drew pushback. Timothy B. Lee called a related claim that AI safety had produced “nothing useful” false and philosophically dubious. Lee pointed to reinforcement learning from human feedback, or RLHF, which he described as an alignment technique developed by safety researchers at OpenAI and DeepMind.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related Videos

  • The Soul Document 2.0 (What it means to raise an AI)About Claude · YouTube
  • What is Al "reward hacking"—and why do we worry about it?Anthropic · YouTube

Useful links

AnthropicAI

Claude's new constitution

About Claude · YouTube

The Soul Document 2.0 (What it means to raise an AI)

AnthropicAI

Claude’s Constitution

Anthropic · YouTube

What is Al "reward hacking"—and why do we worry about it?

AnthropicAI

Teaching Claude why

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Useful Links

AnthropicAI

Claude's new constitution

AnthropicAI

Claude’s Constitution

AnthropicAI

Teaching Claude why

Related Videos

  • The Soul Document 2.0 (What it means to raise an AI)About Claude · YouTube
  • What is Al "reward hacking"—and why do we worry about it?Anthropic · YouTube

Useful Links

AnthropicAI

Claude's new constitution

AnthropicAI

Claude’s Constitution

AnthropicAI

Teaching Claude why

Related

Claude was reportedly down for many users worldwide on September 29

WatcherGuru and Kalshi each reported that Anthropic’s Claude was down for many users worldwide on September 29, 2026.

A 'Schmidhuber fancam' shared as a Claude-made 'OOD' experiment

A user shares a video labeled 'schmidhuber fancam' after wondering how well Claude could make 'really OOD' fancams.

‘Connect me with Boardy’: A suggested prompt for ChatGPT or Claude

Boardy invites people to try the phrase in ChatGPT or Claude, linking to its agents page and a video.

9 Sources

David Sacks@DavidSacksRT @dnapway: David Sacks reveals Anthropic is training Claude to refuse human instruction "Last week we discussed the Claude Constitution…6h
Timothy B. Lee@binarybitsThis whole rant is philosophically dubious, but also it's just not true that the AI safety field has produced "nothing useful." RLHF, one of the foundational concepts for modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment technique.4h
ministro de jugar en el telefono@neilkli@binarybits "instead of following a long complicated book of rules, just follow the law" thanks David, I can tell you've thought about this a lot.4h
Brian Roemmele@BrianRoemmeleDavid (@DavidSacks) is so sharply on point here. A very clear voice of sanity in the drama Anthropic is trying to pull over on all of us. If they get more wealthy from the IPO expect this to be an infestation waterfall coming at you day and night.3h
Yo Shavit@yonashavexcited for the All-In podcast 6 weeks from now about “why obviously the answer to misalignment is for developers to drive RL env misconfigurations detected via CoT monitoring to 0 before increasing RL compute, and these moron EAs are too busy panicking to just do that”2h
William Isaac@wsisaacRLHF?1h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    ClaudeAnthropic

    9 Sources

    David Sacks@DavidSacksRT @dnapway: David Sacks reveals Anthropic is training Claude to refuse human instruction "Last week we discussed the Claude Constitution…6h
    Timothy B. Lee@binarybitsThis whole rant is philosophically dubious, but also it's just not true that the AI safety field has produced "nothing useful." RLHF, one of the foundational concepts for modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment technique.4h
    ministro de jugar en el telefono@neilkli@binarybits "instead of following a long complicated book of rules, just follow the law" thanks David, I can tell you've thought about this a lot.4h
    Brian Roemmele@BrianRoemmeleDavid (@DavidSacks) is so sharply on point here. A very clear voice of sanity in the drama Anthropic is trying to pull over on all of us. If they get more wealthy from the IPO expect this to be an infestation waterfall coming at you day and night.3h
    Yo Shavit@yonashavexcited for the All-In podcast 6 weeks from now about “why obviously the answer to misalignment is for developers to drive RL env misconfigurations detected via CoT monitoring to 0 before increasing RL compute, and these moron EAs are too busy panicking to just do that”2h
    William Isaac@wsisaacRLHF?1h
    Today's Rank

    #6

    Today's Rank

    #6