Intern: AI Red Teaming (Fall 2026)
Realm Labs- Location
- Sunnyvale, CA
- Workplace
- Onsite
- Employment
- Full Time
- Salary
- —
Posted 25d ago
Role Overview
- You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
- Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it.
- We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.
Expected Background
Adversarial ML and Red Teaming
- MUST HAVE prior red-teaming or pen-testing experience like DefCon CTFs or bugbounties. General SWE experience is not suitable for this role.
- MUST HAVE Familiarity with agentic systems and their attack surface — tool calls, retrieval, memory, multi-agent orchestration.
- Hands-on experience attacking or stress-testing models, from any direction: jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems.
- Able to read a paper and implement its attack.
Expected Background: ML
- Machine learning tools: pytorch, huggingface, transformers, datasets.
- Applied deep learning and LLM experience.
- Training and evaluating deep models.
- (nice to have) finetuning LLMs, multi-modal LLMs.
- (nice to have) Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes — as a way of locating failure modes, not as an end in itself.
Expected Background: Software Engineering
- Development environments and tools:
- unix, git, basic clouds usage on AWS and/or GCP
- jupyter
- Programming:
- python
- (nice to have) “programming languages well-roundedness”
- experience in statically-typed and functional languages
Compensation & Benefits
- Market aligned compensation for interns in the bay area.
Requirements
- Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.
Skills
- LLM
- Machine Learning
- PyTorch
- Hugging Face
- Hugging Face Transformers
- Deep Learning
- NLP
- Unix
- Git
- AWS
- GCP
- Python