Headlines about AI safety are usually about hackers. The documented
reality (real incidents, with real victims) is less dramatic and more
useful to know. In rough order of how likely each one is to actually
happen to someone like you:
The AI makes an honest mistake. No attacker
anywhere; it misreads the situation and deletes, renames, or
overwrites the wrong thing. Small versions of this are routine and
Git undoes them. Big versions are rare, but nearly every documented
case of an AI genuinely hurting someone was this: the AI's own error,
reaching further than it should have been able to.
A scam aimed at you, not the AI. Fake
error pages and fake tutorials that talk people into pasting a
"fix-it" command into their own terminal. This is currently the
single most common way real people get hacked, and it targets
exactly the population that follows instructions to paste terminal
commands. That's us.
Poisoned building blocks. Real projects use
lots of small published code packages, and attackers poison some;
hundreds of thousands of malicious ones were logged last year,
including poison hidden inside genuinely popular, normal-looking
packages. "Being careful" doesn't defend you here; nobody can spot
these by reading.
Hidden instructions steering the AI (Prompt
injection). Anything
Claude reads (a webpage, a downloaded file, a stranger's
code) can contain hidden instructions, and Claude can't always tell
"content I'm reading" apart from "orders I should follow". This is
the famous risk, and today it's also the rarest for hobby projects,
but it can't be fixed, only contained, which is why it shapes the
table further down.
The one idea to hold onto
Safety isn't one magic setting. It's two layers: walls
(limits on what the AI can touch, so mistakes and attacks stay small)
and habits (the rules below). The walls catch what
your judgment misses; the habits cover what no wall can reach.
What your cage protects you from.
Most learners run the tutor inside the cage: a
sealed-off world inside your computer. From inside it, the rest
of your machine doesn't exist. Not hidden, not locked: absent.
That one fact does a lot of quiet work:
Your files can't be touched. Photos, documents,
emails, browser history: if Claude goes looking, there is nothing
there to find. Even a fully confused (or fully tricked) Claude can't
reach what isn't there.
There's almost nothing to steal. The cage holds
your project files and one GitHub sign-in token. No saved passwords,
no bank details, no tax files.
Poison lands in an empty room. If a bad package
ever runs inside the cage, it finds lesson files and that one token,
and flt reset rebuilds the whole world, clean, with your
projects kept.
Your home network is blocked. Printers, other
computers, smart-home gadgets: the cage can reach the open internet
but not the devices in your house.
Two honest gaps, so you know exactly what the cage does
not do:
Because the open internet is allowed (you need it to learn), a
tricked Claude could still send the project folder's
contents somewhere. The cage limits what there is to steal, not the
sending. That's why anything private stays out of the cage.
The cage can't protect you from scams where you do the
action, like pasting a command from a fake error page, or typing a
password where you shouldn't. No wall stands between your own hands
and your own machine. The rules below are the defence for that.
A very important note
The cage isn't 100% unbreakable, and you should keep that in mind.
If something seems off, don't assume everything is ok because you
have a cage. Don't encourage Claude to break the cage; instead
reach out to us if you have problems with the cage.
Permission modes: choosing who watches.
Claude Code has "permission modes": it can stop and ask you before
actions, run with a background safety-checker, or run with no checks
at all. Which one is right comes down to a single question:
is there a wall around Claude, or not?
Inside the cage: bypass mode.
Your tutor runs Claude in bypass mode: no
popups, and no background checker either. That sounds alarming, so
here's the reasoning. Anthropic's own guidance says bypass mode
belongs only inside a sealed-off environment, and the cage is
exactly that: it was built so that nothing Claude does can reach the
rest of your computer. With the wall doing the protecting, a checker
on top mostly gets in the way: the actions it blocks are often the
very rescue commands you need when something's gone wrong ("undo my
changes", "delete the broken version"), and its "no" can be silent,
confusing, and impossible to argue with. We'd rather you were never
stuck. Inside the cage, your review moment is the tutor's
pause-before-changes in the chat. That habit, not a popup, is what
keeps you in charge.
One honest limit: the cage's wall protects
the rest of your computer, not your projects; everything
you've built lives in the FLT folder, inside the wall. Your
protection there is the save-point habit: your tutor makes git
save-points from day one precisely so that anything lost can be
brought back.
Outside the cage: auto mode.
Anywhere there's no cage (the cageless setup, or any Claude
you ever run outside the walls), the mode to use is
auto mode. No popup flood here either; instead a
second AI quietly checks each thing Claude wants to
do, in the background, and interrupts only when something looks
dangerous. It's built to catch exactly the accidents that top the
"what actually goes wrong" list: mass deletions, destructive Git
commands, "download this and run it".
What no checker can do: it's judgment, not a
wall. Things slip past it precisely when they look routine:
An "almost right" edit that quietly breaks something looks like a
normal edit.
Installing a poisoned package looks exactly like installing a
normal one, because it is a normal install of a package
someone poisoned upstream.
Publishing code with a secret key accidentally pasted inside
looks like any other publish.
And scams aimed at you never pass through Claude at all, so the
checker never even sees them.
That's why the careful habits on this page matter everywhere, in
every mode; no watcher replaces them.
Bypass belongs inside walls
The rule of thumb: bypass mode only ever inside the
cage. Out there (no wall, no checker) bypass mode means
nothing is watching at all, and one confused delete can reach
anything on the machine. If anything (a tutorial, a forum post, an
AI) suggests running Claude with "bypass permissions"
outside a sealed-off environment, the answer is no.
Four rules that work everywhere.
These apply in the cage, out of the cage, on the Claude website,
everywhere. Between them they cover the doors no wall can close.
Commands and downloads come only from your tutor or this
site. Never from a pop-up, an error page, a forum, or a
stranger's video. The "paste this command to fix it" scam is the
single most common way real people get hacked, and it works by
getting you to do the pasting.
Never type a real password into anything
Claude-adjacent. Not into the terminal, not into a chat.
Signing in happens in your browser, on the real website, and nowhere
else.
Remember: anything Claude reads can steer what Claude
does. So the question "where should this project live?" is
really "will Claude be reading things I didn't write and don't
trust?", and if yes, it belongs somewhere Claude has walls, or no
hands at all. The table below does this sorting for you.
Protect the accounts themselves. Your GitHub and
Claude accounts are part of your safety. Two cheap habits defeat
almost all real-world account attacks: a password you use
nowhere else, and two-step sign-in (where the site also
checks your phone) turned on.
Which project needs which safety?
We decide by asking a question: how risky is this
thing?
Note: Risk lives in what a project reads, runs and touches, not
in the person doing it. A stranger's webpage is exactly as dangerous
for a professional as for a beginner.
CLAUDE WEBSITE A plain chat
at claude.ai: no files, no
commands. A steered Claude here has no hands; the worst case is a
wrong answer. ("Plain" matters: no connected email or drive, no
personal details mixed into a chat that's reading strangers'
content.)
CAGE The walled world you
already have. Most projects live here, many forever.
OUTSIDE THE CAGE An
ordinary Claude Code install with no cage, for the few projects
that genuinely wouldn't work inside it, because they need to touch
things the cage deliberately blocks. Claude has its own built-in
sandbox setting (a lighter set of walls we'd help you set up when
the time comes), and the rest is careful habits: auto mode on, Git
religiously, smallest possible scope. Not part of the course; a
later, deliberate step.
What you're doing
How risky?
Why it's ranked that way
Best home
The course itself
LOW
Tiny exercises with nothing at stake, and mistakes are
part of the point.
CAGE
A personal website, portfolio, or browser game, published
via GitHub
LOW
Public by design, nothing private involved. Worst case is
a broken page, and Git undoes it.
CAGE
Reading things from the internet for you: "summarise this
article", "compare these two products"
MEDIUM
Web pages you didn't write can carry hidden instructions
that try to steer the AI.
CLAUDE WEBSITE: in a
plain chat, the worst a steered AI can do is give you a bad
answer
Home hardware: a Raspberry Pi, smart-home tinkering
LOW
Stakes are naturally small: mistakes affect a hobby gadget,
not your machine.
OUTSIDE THE CAGE: it
needs your home network, which the cage deliberately blocks
Worth noticing: the two columns answer different questions, and
they don't always move together. The Raspberry Pi row is
low risk yet lives outside the cage, because the
home is decided by what a project needs to reach, and a Pi needs your
home network, which the cage blocks. Risk tells you how careful
the setup has to be; what the project needs to touch tells you which
setup that is.
The day you get your first API key.
Sooner or later a project will want an API key: a
long secret code that lets your project use a paid service, where every
use costs money on your card. That day, four things:
Best trick: keep the key off your machine
entirely. For anything hosted, like a chatbot on your
website, build in the cage with a fake placeholder key, and paste
the real key into your hosting service's "secrets" store, in your
browser. The server holds it; your laptop never does. This is how
professionals ship it.
Set a spending limit before the key does
anything. Every provider offers one.
Keys never go in the project folder. The
classic accident is a key pasted into code and published to GitHub;
millions of secrets leak this way every year, and AI-assisted
projects leak them at roughly double the normal rate.
Revoke is the undo button. Find the "revoke
key" page on the provider's site before you need it. A leaked key
stops mattering the moment it's revoked.
Unsure where a project belongs?
Research first: the WhatsApp group while you're with us, someone
experienced, or a plain chat with your AI. Ask before you start,
not after.