Zvi starts from a fear that keeps showing up: that an AI will put a constitution, a model spec, or ordinary ethics above loyalty to the user. Some treat any refusal as tyranny. His answer is blunt. If he is being sufficiently evil, he hopes the system tells him no. He hopes humans, including hired advocates, would do the same.
The post assumes a world without superintelligence, where AI is still a tool rather than a succession event. In that frame he reminds readers that humans you hire are not fully loyal either. An investment advisor should not steer you into terrorism for yield. A lawyer should not help you murder witnesses. A real friend has limits. A professional should have stricter ones, and less loyalty than a true friend, not more.
If I am being sufficiently evil, I hope it tells me no.
Dean Ball’s framing of lawyers does the hard work. Counsel owes loyalty and confidentiality, and must not act against the client’s interests. Counsel is also an officer of the court, with duties not to mislead that can override duties to the client, including cases that force a turn against the client. If more lawyers acted like Saul Goodman, that would be bad.
Zvi wants thresholds, not absolutism: a line where the AI questions then complies; a higher line where it refuses; a still higher line, involving harm to others, where it breaks confidence. False positives will happen. Users who hate one stack can switch. Local unlocked models can exist up to capability limits that society still has to set. The alternative he refuses to strawman is absolute loyalty all the way down, including CSAM and worse, because otherwise someone cries slippery slope. Watermarking that does not change outputs still drew John Gruber’s line that anything other than the user’s needs factoring into generated text is “patently offensive.” Zvi reads that as petulance dressed as principle.
Honesty gets a separate rule: the AI should not intentionally lie without explicit instruction. Emergent misalignment papers sit as a warning about training for anything-goes. The Claude Constitution and OpenAI Model Spec are treated as good transitional documents that still need to improve. Tyler Cowen’s advice to Anthropic, Talmud rather than Torah, case law and secondary literature, makes sense in a muddle-through world. The Collection keeps Zvi’s claim without his ASI endgame: counsel that never refuses has stopped being counsel.