Security Of Enterprise Agents, Frontier Labs Challenge, Superworker Wages
Today I review the various jailbreaks by both OpenAI and Anthropic and the challenges which Frontier Labs are having with AI “alignment.” So in keeping with HR 2030, I share some general advice and stay tuned for our detailed technical paper on this in the coming weeks.
And as we debate the reliability, training, and alignment of models I also point out how the technology industry is struggling to deal with the politics and economics here, since right now most tech leaders really don’t want any form of regulation. (Unless it’s somehow fair to all.) So we, as buyers, are left dealing with the issue.
AI Alignment means making sure an AI system’s goals and behavior match what people actually want—our values, rules, and intentions. It’s about getting the AI to do the “right thing” even in new situations, not just follow instructions literally in ways that cause harm. In practice, it includes preventing unwanted outcomes like deception, unsafe shortcuts, or optimizing a metric that misses the real objective.
I also highly recommend Jill Lepore’s book “The Rise of the Artificial State” for a very sound discussion of the political issues around AI, which helps with these issues.
And along the line of AI “alignment,” I discuss why Google paid $10M for the emails and data from bankrupt Spirit Airlines, and how AI companies are now buying old slack and emails from many bankrupt companies, training their AI on old, failed business models. (Read below on “reinforcement learning gyms.”)
This is a strange series of events, all making it clear that Google, OpenAI, and Anthropic have quite a challenge training and aligning their models in the future.
And I conclude with a preview of our AI-driven wage research and our coming book Superpowered, coming in October.
Resources You Should Read
OpenAI cyber models broke out of training environment to hack Hugging Face
Anthropic’s disclosure of Claude’s jailbreak
Google to Buy Spirit Airlines Business Data for $10 Million
A.I.’s New Training Data: Your Old Work Slacks and Emails
The Birth of “Reinforcement Learning Gyms”
Cisco Gave All 90,000 Employees Their Own AI Agent
A Peek Into The Workforce Of 2035: Jobs, Wages, And AI’s Impact (Bersin)
CEO Pay Ratios (vs. average employee) for Fortune 500