WebsiteHunt is broughtBrought to you by Redact Everything
SS

Small-Samples Poison in LLMs

Anthropic research on data-poisoning attacks in LLMs.

This page summarizes Anthropic's research on data-poisoning in large language models. The study shows that as few as 250 malicious documents can backdoor models from 600M to 13B parameters, challenging the idea that poisoning scales with data volume. It discusses backdoors, risks, and directions for mitigation and defense.

Comments

Loading comments...

You might also like

HiAPI

HiAPI

One API for leading AI image, video, and text generation models. ✨ Premium
ProjectManagementTools

ProjectManagementTools

Optimize your Projects with ProjectManagementTools 🌱 Sponsored
Labrynth AI

Labrynth AI

AI regulatory intelligence platform 🌱 Sponsored
Anthropic: AI Safety & Policy

Anthropic: AI Safety & Policy

Responsible AI research and policy updates from Anthropic.
Jina Reader API

Jina Reader API

Read URLs and search web for better grounding LLMs.
Frontiers in Political Science

Frontiers in Political Science

Open-access, peer-reviewed research in political science.
Buzz App

Buzz App

AI-powered Market Research
Introducing Claude Opus 5

Introducing Claude Opus 5

Anthropic's model for long-running agents and coding.
Position: LLMs can't jump

Position: LLMs can't jump

A position paper on LLMs, abduction, and scientific invention.
Stantem

Stantem

Property Search, Real Estate Data API & Data Aggregation Services
NH

New Hormone Builds Strong Bones

Researchers uncover a hormone strengthening bones in lactation.
Wiz Blog

Wiz Blog

Security research and vulnerability disclosures in AI.
PlasticList

PlasticList

Data on plastic chemicals in Bay Area foods
Korvus

Korvus

Search SDK to unify RAG pipeline in a single database query
ChatPulse

ChatPulse

Gain real insights on team communication by leveraging Slack data
RamenLegal

RamenLegal

AI suite for legal documentation and research.
TimeCapsuleLLM

TimeCapsuleLLM

An LLM trained only on era-specific data to reduce modern bias.
Elicit

Elicit

The AI research assistant
Juno Labs Blog

Juno Labs Blog

Thoughtful AI insights on privacy, data, and policy.
StratChat Premium

StratChat Premium

An AI-driven search engine for stocks, and investing
DataFuel.dev

DataFuel.dev

Turn websites into LLM-ready data.
Keyword Insights

Keyword Insights

Save hours doing keyword research
GG

GreyNoise Grimoire

Security research and insights into Internet background noise.
by @Micadep