programming.dev
  • Communities
  • Create Post
  • Create Community
  • heart
    Support Lemmy
  • search
    Search
  • Login
  • Sign Up
ProM to AI - Artificial intelligenceEnglish · 1 month ago

Research leaders urge tech industry to monitor AI’s ‘thoughts’

www.greaterwrong.com

external-link
message-square
0
link
fedilink
0
external-link

Research leaders urge tech industry to monitor AI’s ‘thoughts’

www.greaterwrong.com

ProM to AI - Artificial intelligenceEnglish · 1 month ago
message-square
0
link
fedilink
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
www.greaterwrong.com
external-link
Twitter | Paper PDF Seven years ago, OpenAI five had just been released, and many people in the AI safety community expected AIs to be opaque RL agents. Luckily, we ended up with reasoning models that speak their thoughts clearly enough for us to follow along (most of the time). In a new multi-org position paper, we argue that we should try to preserve this level of reasoning transparency and turn chain of thought monitorability into a systematic AI safety agenda. This is a measure that improves safety in the medium term, and it might not scale to superintelligence even if somehow a superintelligent AI still does its reasoning in English. We hope that extending the time when chains of thought are monitorable will help us do more science on capable models, practice more safety techniques "at an easier difficulty", and allow us to extract more useful work from potentially misaligned models.
alert-triangle
You must log in or register to comment.

AI - Artificial intelligence

Aii

Subscribe from Remote Instance

Create a post
You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: [email protected]

AI related news and articles.

Rules:

  • No Videos.
  • No self promotion: Don’t post links to your articles.
Visibility: Public
globe

This community can be federated to other instances and be posted/commented in by their users.

  • 22 users / day
  • 45 users / week
  • 375 users / month
  • 503 users / 6 months
  • 16 local subscribers
  • 85 subscribers
  • 150 Posts
  • 72 Comments
  • Modlog
  • mods:
  • Pro
  • BE: 0.19.11
  • Modlog
  • Legal
  • Instances
  • Docs
  • Code
  • join-lemmy.org