Research & safety

Follow what AI can do—and where it can fail

Track capability research alongside the evidence needed to judge reliability, security, privacy, alignment, misuse and frontier-model risk.

Research frontier

Capabilities, agents and scientific AI

Track research that changes what AI systems can do: reasoning, multimodality, tool use, agents, long context, coding, mathematics and scientific work.

reasoningagentsmultimodal AIcodingmathematicssciencelong contexttool useruntime verification
Testing

Evaluations, red teaming and model assurance

Model evaluations are useful only when the test actually matches the claim. Follow benchmark design, red teaming, cyber testing, external evaluations and evidence limitations.

benchmarksred teamingexternal evaluationcyber testsrobustnessassurancemeasurement limitstrajectory verification
Frontier safety

Alignment, control and high-capability model risk

Follow research on keeping advanced systems controllable, robust and aligned with intended goals as autonomy and capability increase.

alignmentcontrolautonomyagent safetyrobustnessoversightfrontier modelsruntime contracts
Security

Cybersecurity, prompt injection and agent permissions

AI systems can create new attack surfaces when they can browse, call tools, access credentials or take actions. Security coverage focuses on permissions, containment and recovery.

prompt injectiontool permissionscredentialssandboxingcyber capabilityleast privilegeaudit trailsproof of execution
Trust

Hallucinations, factuality and reliability

AI can produce fluent answers without reliable evidence. Track hallucination research, verification methods, uncertainty, source use and the limits of detection systems.

hallucinationsfactualityuncertaintycitationsverificationAI detectionsource qualitycompletion evidence
Data protection

Privacy, sensitive data and information leakage

AI privacy risk depends on what data is shared, how it is retained, who can access it and whether the system is allowed to call external tools or services.

privacydata retentionsensitive datadata residencyinformation leakagelocal AIaccess control
Misuse

Deepfakes, fraud, synthetic media and abuse

Track harmful uses of generative AI and the evidence behind proposed defenses, including provenance, labeling, detection and platform controls.

deepfakesfraudsynthetic mediamisinformationprovenancecontent labelsdetection