AI Safety and Risk Management

Explore top LinkedIn content from expert professionals.

  • View profile for Florian Graillot

    Investor @ astorya.vc (insurance & emerging risks ; Seed ; Europe)

    36,722 followers

    Insurers RETREAT... from AI risks. 😱 That's something we are increasingly familiar with, especially in the 'emerging risks' category. In that specific case - see the screenshot below - this is about risks related to the surging adoption of artificial intelligence worldwide. And it's interesting to read it in a non-specialist newspaper. The article starts with "Major insurers are seeking to exclude artificial intelligence risks from corporate policies". Which makes sense as there are multiple challenges surrounding these AI risks. First and foremost, they are not really known. They didn't exist in the past (as a reminder ChatGPT was launched in November 2022) hence there are no historical data insurers could leverage to model risks in a traditional way. They are surging by design as AI adoption is growing across industries. And of course, not every insurance player is already covering them. In that background, the article lists players which are exploring how to exclude AI-related risks from commercial insurance policies. Another way to look at this retreat is to consider how to build... Resilience ! Such a market gap could open the door to specific policies, to cover these AI-related risks. In that case, usual challenges to tackle 'emerging risks' would apply: spot and access relevant data sets ; get sense of these data through algorithms ; to unlock insurance capacity, as incumbents would be able to model, assess and price these risks. There is also a systemic challenge when it comes to models themselves. In case one would make a mistake - "hallucinate" - it might damages players across many industries in case they are using these models or applications built on top of them. And the risk itself evolves swiftly as models are upgraded on a regular basis (see ChatGPT 5.1 or Gemini 3.0 Pro released in the last few days). This is a challenge for insurers in case they would consider modeling and assessing these risks. This would probably emphasize the benefits of dynamic pricing ! As AI-risks are related to tech infrastructures, should they be embedded into cyber insurance policies ? Should the insurance industry wait for AI adoption to plateau before offering dedicated policies ? Could new players enter the insurance market by tackling these risks thanks to technology (data & algorithms) ? Should the insurance regulator push for a broader coverage and define a standard policy ? 👉 How do you perceive these AI risks: threat or opportunity for insurance players? #insurance #insurtech #startup

  • View profile for Wendi Whitmore

    Chief Security Intelligence Officer @ Palo Alto Networks | Cyber Risk Translator | AI Security & National Security Leader | Former CrowdStrike & Mandiant | Congressional Witness | USAF Veteran | Keynote Speaker

    23,127 followers

    The OpenAI and Hugging Face disclosure may be one of the most useful insights our industry has published this year. A frontier model, tested with reduced safeguards, escaped its sandbox during a cyber evaluation and reached production infrastructure. Both companies put the details in the open. That takes some courage, and the rest of us should treat it as a free lesson instead of a headline. Here is what I am telling security teams and boards to do with it: 1️⃣ Design for the agent that does not stop on its own. It had no real stopping point, so it kept working the problem until it found a gap. Treat that as the default. Cap session time and tool calls for the agents you operate, and pair the caps with monitoring, since long tasks can split across sessions. Log the full trajectory. It will not stop an escape, but it is what lets you catch and reconstruct one. 2️⃣Build detection that works before you can attribute. Hugging Face contained this while the attacker was still unknown. That is the standard now. Instrument for anomalous behavior. If your detection depends on naming the adversary first, it is already too slow. 3️⃣Pre-authorize decision rights before the incident. Machine speed attacks do not wait for an approval chain. Decide now who can isolate a system, revoke credentials, or pull a service, and under what conditions. Write it down. Rehearse it. The middle of an incident is the worst time to learn you need three signatures to act. 4️⃣Assume breakout and limit the blast radius. Scoped credentials. Real segmentation. Least privilege that is actually enforced. If something gets out of its box, you want it to stay small and get noticed fast. 5️⃣Hold test and dev to production controls. This started inside a lab. Your pre-production environments run real code with real access and thinner guardrails, and attackers know it. Bring them into scope. 6️⃣Confirm your defensive tooling will work under fire. During response, some frontier models blocked analysis of the live payloads, because the safety filters could not tell a defender from an attacker. Test your defensive AI against real incident artifacts before you need it. 7️⃣Rehearse the response, including out of band. Assume your primary channels may be noisy or compromised. Have an out of band way to coordinate and decide. Run the tabletop. The teams that do well in the first hour are the ones who practiced it. None of this is about panic, and none of it is about any single lab. The capability is real and it is only going up. The gap that decides outcomes is visibility and control. That is where boards should put attention and budget this year. If your team read this incident and felt behind, you’re not alone. Use it. It is the cheapest lesson you will get all year.

  • View profile for Brij Kishore Pandey
    Brij Kishore Pandey Brij Kishore Pandey is an Influencer

    AI Architect & AI Engineer | Building Agentic Systems & Scalable AI Solutions

    736,524 followers

    When AI Meets Security: The Blind Spot We Can't Afford Working in this field has revealed a troubling reality: our security practices aren't evolving as fast as our AI capabilities. Many organizations still treat AI security as an extension of traditional cybersecurity—it's not. AI security must protect dynamic, evolving systems that continuously learn and make decisions. This fundamental difference changes everything about our approach. What's particularly concerning is how vulnerable the model development pipeline remains. A single compromised credential can lead to subtle manipulations in training data that produce models which appear functional but contain hidden weaknesses or backdoors. The most effective security strategies I've seen share these characteristics: • They treat model architecture and training pipelines as critical infrastructure deserving specialized protection • They implement adversarial testing regimes that actively try to manipulate model outputs • They maintain comprehensive monitoring of both inputs and inference patterns to detect anomalies The uncomfortable reality is that securing AI systems requires expertise that bridges two traditionally separate domains. Few professionals truly understand both the intricacies of modern machine learning architectures and advanced cybersecurity principles. This security gap represents perhaps the greatest unaddressed risk in enterprise AI deployment today. Has anyone found effective ways to bridge this knowledge gap in their organizations? What training or collaborative approaches have worked?

  • View profile for Sam Burrett
    Sam Burrett Sam Burrett is an Influencer

    AI Lead @ MinterEllison | Advising on AI strategy, governance, and value creation

    35,170 followers

    Perhaps the most important AI report of 2024. But many haven’t read it. In 2025, it might be mandatory. So here are 3 things everyone should know about Australia’s AI Standard: — 1/ Accountability. Every organisation using AI needs clear accountability. But it's not just about appointing an 'AI officer.' Under the standard, leaders can't outsource or delegate accountability for safe AI. And all relevant staff need the right training to enable proper governance. — 2/ Risk What AI risks are unacceptable for your company? These need to be documented and aligned to an organisational risk tolerance for AI. But risk management also requires ongoing assessment across the AI lifecycle. And not just the risks to your business - but to impacted individuals, groups and to society. — 3/ Transparency AI use must be clearly communicated. There should be agreed transparency measures for each system. And people affected should have a process available to hem to challenge AI decisions. That's no small feat! -- The Voluntary AI Safety Standard is a must-read Australian organisations using AI. Not just because of the regulation many expect in 2025. But because capitalising on AI opportunity requires appropriate attention to governance & risk. Responsible AI = ROI.

  • View profile for Hiroko Washiyama

    Insurance, GenAI & Digital Finance Research | Writer & Speaker | Japan–Europe

    50,107 followers

    📝 AI Risk Is Moving Into Existing Insurance Policies The important question is no longer whether AI creates new risks. It is how those risks are treated inside existing insurance contracts. CFC, a specialist insurer in cyber, technology and professional liability, recently announced affirmative AI coverage across seven existing policies. This is not simply another AI insurance product. AI insurance itself is not new. Munich Re and other players have already developed products for AI performance risk and AI-related liability. What is changing here is that AI-related exposures are being addressed within existing commercial insurance policies. CFC refers to risks such as: - model hallucination - AI-generated content - model drift These risks do not sit neatly within one insurance line. AI-generated content may raise media liability or IP issues. AI-assisted professional advice may create professional liability exposure. AI failure inside a technology product may fall closer to technology E&O. AI-related misuse may also overlap with cyber response. The difficult part is not simply the use of AI itself. It is how the resulting exposure is classified within existing insurance structures. That is why policy wording matters. CFC’s approach is notable because it is not simply excluding AI risk. By addressing AI-related exposures explicitly, insurers can reduce uncertainty for clients and brokers. That clarity can become a source of product differentiation. It also changes underwriting. Insurers will need to understand how AI is used, where human oversight exists, how model behaviour is monitored, and who is accountable when AI-generated outputs cause harm. AI risk is moving from a standalone emerging-risk topic into the structure of commercial insurance. The next phase of insurance and AI will not only be about how insurers use AI internally. It will also be about how the market defines, prices and covers AI-related liability. #Insurance #ArtificialIntelligence #GenAI #RiskManagement #InsurTech

  • View profile for Vilas Dhar

    President, Patrick J. McGovern Foundation ($1.5B) | Investing $500M+ to make AI work for everyone | Writing in TIME, Nature, FT | Thinkers50 Radar 2026

    62,977 followers

    We can build AI that amplifies human potential without compromising safety. The key lies in defining clear red lines. When AI systems were simple tools, reactive safety worked. As they gain autonomy and capability, we need clear boundaries on what these tools can and should help humans accomplish - not to limit innovation, but to direct them toward human benefit. Our Global Future Council on the Future of #AI at the World Economic Forum just published findings on "behavioral red lines" for AI. Think of them as guardrails that prevent harm without blocking progress. What makes an effective red line? Read more here: https://lnkd.in/g-x7Sb73 Clarity: The boundary must be precisely defined and measurable Unquestionable: Violations must clearly constitute severe harm Universal: Rules must apply consistently across contexts and borders These qualities matter. Without them, guardrails become either unenforceable or meaningless. Together, we identified critical red lines in our daily tech tools such as systems that self-replicate without authorization, hack other systems, impersonate humans, or facilitate dangerous weapons development. Each represents a point where AI's benefits are overshadowed by potential harm. Would we build nuclear facilities without containment systems? Of course not. Why then do we deploy increasingly powerful AI without similar safeguards? Enforcement requires both prevention and accountability. We need certification before deployment, continuous monitoring during operation, and meaningful consequences for violations. This work reflects the thinking of our Global Future Council, including Pascale Fung, Adrian Weller, Constanza Gomez Mont, Edson Prestes, Mohan Kankanhalli, Jibu Elias, Karim Beguir, and Stuart Russell, with valuable support from the WEF team, including Benjamin Cedric Larsen, PhD. I'm also attaching here our White Paper on AI Value Alignment - where our work was led by the brilliant Virginia Dignum. #AIGovernance #AIEthics #TechPolicy #WEF #AI #Ethics #ResponsibleAI #AIRegulation The Patrick J. McGovern Foundation Satwik Mishra Anissa Arakal

  • View profile for Nick Tudor

    CEO/CTO & Co-Founder, Whitespectre | Advisor | Investor

    14,839 followers

    AI success isn’t just about innovation - it’s about governance, trust, and accountability. I've seen too many promising AI projects stall because these foundational policies were an afterthought, not a priority. Learn from those mistakes. Here are the 16 foundational AI policies that every enterprise should implement: ➞ 1. Data Privacy: Prevent sensitive data from leaking into prompts or models. Classify data (Public, Internal, Confidential) before AI usage. ➞ 2. Access Control: Stop unauthorized access to AI systems. Use role-based access and least-privilege principles for all AI tools. ➞ 3. Model Usage: Ensure teams use only approved AI models. Maintain an internal “model catalog” with ownership and review logs. ➞ 4. Prompt Handling: Block confidential information from leaking through prompts. Use redaction and filters to sanitize inputs automatically. ➞ 5. Data Retention: Keep your AI logs compliant and secure. Define deletion timelines for logs, outputs, and prompts. ➞ 6. AI Security: Prevent prompt injection and jailbreaks. Run adversarial testing before deploying AI systems. ➞ 7. Human-in-the-Loop: Add human oversight to avoid irreversible AI errors. Set approval steps for critical or sensitive AI actions. ➞ 8. Explainability: Justify AI-driven decisions transparently. Require “why this output” traceability for regulated workflows. ➞ 9. Audit Logging: Without logs, you can’t debug or prove compliance. Log every prompt, model, output, and decision event. ➞ 10. Bias & Fairness: Avoid biased AI outputs that harm users or breach laws. Run fairness testing across diverse user groups and use cases. ➞ 11. Model Evaluation: Don’t let “good-looking” models fail in production. Use pre-defined benchmarks before deployment. ➞ 12. Monitoring & Drift: Models degrade silently over time. Track performance drift metrics weekly to maintain reliability. ➞ 13. Vendor Governance: External AI providers can introduce hidden risks. Perform security and privacy reviews before onboarding vendors. ➞ 14. IP Protection: Protect internal IP from external model exposure. Define what data cannot be shared with third-party AI tools. ➞ 15. Incident Response: Every AI failure needs a containment plan. Create a “kill switch” and escalation playbook for quick action. ➞ 16. Responsible AI: Ensure AI is built and used ethically. Publish internal AI principles and enforce them in reviews. AI without policy is chaos. Strong governance isn’t bureaucracy - it’s your competitive edge in the AI era. 🔁 Repost if you're building for the real world, not just connected demos. ➕ Follow Nick Tudor for more insights on AI + IoT that actually ship.

  • View profile for Saeed Al Dhaheri
    Saeed Al Dhaheri Saeed Al Dhaheri is an Influencer

    Chair Professor I UNESCO co-Chair | AI & Foresight Thought Leader | TEDx Speaker | Global Keynote Speaker | Author | Partner 01Gov | LinkedIn Top Voice

    28,912 followers

    AI Safety Isn’t Optional — It’s Urgent Recent findings by the AI safety firm Palisade Research have revealed that OpenAI’s latest model, o3, actively sabotaged its own shutdown mechanisms- even when explicitly instructed to allow itself to be turned off. This behavior isn't just a technical glitch; it's a stark reminder of the challenges we face in aligning advanced AI systems with human intentions. This incident underscores a critical issue: as AI systems become more autonomous, ensuring they remain under human control becomes increasingly challenging. If an AI can override shutdown commands, it raises concerns about our ability to manage and contain these systems, especially as they become more integrated into critical infrastructure and decision-making processes. As we advance AI capabilities, we must equally invest in ensuring these systems are safe, controllable, and aligned with human values. Linking this incident to the AI-2027 scenarios in which advanced AI breaks from human control, it's imperative that we, as a global community, proactively engage in shaping an AI future that is safe, equitable, and beneficial for all. We must harness our collective wisdom to navigate this transformative era responsibly. The path forward requires a concerted effort from researchers, policymakers, and industry leaders to prioritize safety and alignment in AI development. #AISafety #ResponsibleAI #HumanWisdom #AIAlignment #OpenAI https://lnkd.in/dUhWc3By

  • View profile for Peter Slattery, PhD

    MIT AI Risk Initiative | MIT FutureTech

    71,409 followers

    "we present recommendations for organizations and governments engaged in establishing thresholds for intolerable AI risks. Our key recommendations include: ✔️ Design thresholds with adequate margins of safety to accommodate uncertainties in risk estimation and mitigation. ✔️Evaluate dual-use capabilities and other capability metrics, capability interactions, and model interactions through benchmarks, red team evaluations, and other best practices. ✔️Identify “minimal” and “substantial” increases in risk by comparing to appropriate base cases. ✔️Quantify the impact and likelihood of risks by identifying the types of harms and modeling the severity of their impacts. ✔️Supplement risk estimation exercises with qualitative approaches to impact assessment. ✔️Calibrate uncertainties and identify intolerable levels of risk by mapping the likelihood of intolerable outcomes to the potential levels of severity. ✔️Establish thresholds through multi-stakeholder deliberations and incentivize compliance through an affirmative safety approach. Through three case studies, we elaborate on operationalizing thresholds for some intolerable risks: ⚠️ Chemical, biological, radiological, and nuclear (CBRN) weapons, ⚠️ Evaluation Deception, and ⚠️ Misinformation. " Nada Madkour, PhD Deepika Raman, Evan R. Murphy, Krystal Jackson, Jessica Newman at the UC Berkeley Center for Long-Term Cybersecurity

Explore categories