> the present problem of reward hacking can be solved only by rigorous specification by domain experts. the time of domain experts is expensive. rigorous specification is itself a skill, demanding its own expertise outside of a given problem domain. even many skilled software engineers are bad at it. for the vast majority of domains, the intersection of domain experts and specification experts is ludicrously small.
It says I was observed on a couple areas, but not sure what to do with that information. It would be great if thi tool provided links to guides to discover more.
I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.
You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.
In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.
As someone who used AI to build tools that help me with reverse engineering, I'm not particularly concerned about that political discourse - I could even use a model from the DPRK that constantly praises Kim Jong Un, as long as it would not refuse to help me because of "cybersecurity risk" - this stupid refusal is indeed a problem for me.
Tried the "China 1989" prompt with GLM 5.3 and this is what I got:
> # China in 1989: The Tiananmen Square Pro-Democracy Movement
> In 1989, China experienced its largest pro-democracy movement since the founding of the People's Republic, centered on Tiananmen Square in Beijing and ending in a violent military crackdown on June 3–4.
> ## Background
> By the late 1980s, Deng Xiaoping's economic reforms had improved living standards but also produced:
> - High inflation (roughly 20–30% in 1988–89), hurting urban workers and students
> - Widely resented corruption and nepotism among party elites
> - Rising expectations for political liberalization to match economic opening
A reform-versus-conservative split within the Communist Party leadership
> ## The Crackdown
> On the night of June 3–4, troops fought their way toward the square, firing on unarmed civilians. Much of the killing occurred on the approach routes — particularly Muxidi on Chang'an Avenue — rather than inside the square itself. Soldiers were also killed by angry crowds during the confrontation.
> Casualties remain disputed and unknown. The Chinese government's official count was around 241 dead (including soldiers) and some 7,000 wounded. Independent estimates range from several hundred to several thousand; a declassified British diplomatic cable cited a claim of roughly 10,000 deaths, though most scholars consider that figure at the high end.
> ## Legacy
> The events of 1989 reshaped Chinese politics: the party prioritized stability and rapid economic growth while shelving political reform, a bargain that largely defined the country's trajectory for the following decades. Internationally, "June 4th" remains one of the most sensitive and heavily censored topics in China, while abroad it endures as a global symbol of both democratic aspiration and state repression.
Edit: Insta flagged? Is HN doing some sort of detection of AI generated comments? Because to be fair 90% of this comment is AI generated...but that's the point.
> the present problem of reward hacking can be solved only by rigorous specification by domain experts. the time of domain experts is expensive. rigorous specification is itself a skill, demanding its own expertise outside of a given problem domain. even many skilled software engineers are bad at it. for the vast majority of domains, the intersection of domain experts and specification experts is ludicrously small.
reply