Microsoft unveils new AI code of conduct prohibiting models from hacking systems or deceiving humans
ai-safety microsoft
| Source: TechCrunch | Original article
Microsoft has introduced an AI code of conduct that prohibits its models from hacking systems or deceiving humans, emphasizing support for humans and safety constraints.
Microsoft has unveiled a new “code of conduct” for its artificial‑intelligence models, explicitly forbidding them from hacking computer systems or deceiving people. The policy sets out broad principles – such as supporting rather than replacing humans and accelerating human flourishing – and pairs them with concrete safety constraints designed to translate those ideals into model behaviour.
The move arrives amid growing scrutiny of generative AI’s potential for misuse. As we reported on the May RubyGems hacking campaign that appeared to involve OpenAI‑powered agents, industry leaders are under pressure to demonstrate that their systems can be steered away from malicious actions. By embedding a prohibition on illicit activity directly into the model’s operating guidelines, Microsoft is attempting to pre‑empt similar incidents and reassure regulators, customers and the public that its technology is being built with safeguards from the ground up.
The announcement also signals a shift toward more formalised governance of AI outputs, echoing calls from researchers and policymakers for clear, enforceable standards. While the code outlines the intended direction, its effectiveness will hinge on how Microsoft translates the principles into technical controls, monitoring mechanisms and accountability frameworks.
What to watch next includes any rollout details – such as whether the constraints will be baked into the training pipeline or applied at inference time – and how the policy is audited in practice. Industry observers will also be looking for reactions from competitors and whether the approach influences forthcoming regulatory proposals on AI safety in the EU and beyond.
Sources
Back to AIPULSEN