Staying safe with AI.

What actually goes wrong, what protects you, and where each kind of project should live.

On this page.

What actually goes wrong.

Headlines about AI safety are usually about hackers. The documented reality (real incidents, with real victims) is less dramatic and more useful to know. In rough order of how likely each one is to actually happen to someone like you:

The one idea to hold onto Safety isn't one magic setting. It's two layers: walls (limits on what the AI can touch, so mistakes and attacks stay small) and habits (the rules below). The walls catch what your judgment misses; the habits cover what no wall can reach.

What your cage protects you from.

Most learners run the tutor inside the cage: a sealed-off world inside your computer. From inside it, the rest of your machine doesn't exist. Not hidden, not locked: absent. That one fact does a lot of quiet work:

Two honest gaps, so you know exactly what the cage does not do:

A very important note The cage isn't 100% unbreakable, and you should keep that in mind. If something seems off, don't assume everything is ok because you have a cage. Don't encourage Claude to break the cage; instead reach out to us if you have problems with the cage.

Permission modes: choosing who watches.

Claude Code has "permission modes": it can stop and ask you before actions, run with a background safety-checker, or run with no checks at all. Which one is right comes down to a single question: is there a wall around Claude, or not?

Inside the cage: bypass mode.

Your tutor runs Claude in bypass mode: no popups, and no background checker either. That sounds alarming, so here's the reasoning. Anthropic's own guidance says bypass mode belongs only inside a sealed-off environment, and the cage is exactly that: it was built so that nothing Claude does can reach the rest of your computer. With the wall doing the protecting, a checker on top mostly gets in the way: the actions it blocks are often the very rescue commands you need when something's gone wrong ("undo my changes", "delete the broken version"), and its "no" can be silent, confusing, and impossible to argue with. We'd rather you were never stuck. Inside the cage, your review moment is the tutor's pause-before-changes in the chat. That habit, not a popup, is what keeps you in charge.

One honest limit: the cage's wall protects the rest of your computer, not your projects; everything you've built lives in the FLT folder, inside the wall. Your protection there is the save-point habit: your tutor makes git save-points from day one precisely so that anything lost can be brought back.

Outside the cage: auto mode.

Anywhere there's no cage (the cageless setup, or any Claude you ever run outside the walls), the mode to use is auto mode. No popup flood here either; instead a second AI quietly checks each thing Claude wants to do, in the background, and interrupts only when something looks dangerous. It's built to catch exactly the accidents that top the "what actually goes wrong" list: mass deletions, destructive Git commands, "download this and run it".

What no checker can do: it's judgment, not a wall. Things slip past it precisely when they look routine:

That's why the careful habits on this page matter everywhere, in every mode; no watcher replaces them.

Bypass belongs inside walls The rule of thumb: bypass mode only ever inside the cage. Out there (no wall, no checker) bypass mode means nothing is watching at all, and one confused delete can reach anything on the machine. If anything (a tutorial, a forum post, an AI) suggests running Claude with "bypass permissions" outside a sealed-off environment, the answer is no.

Four rules that work everywhere.

These apply in the cage, out of the cage, on the Claude website, everywhere. Between them they cover the doors no wall can close.

  1. Commands and downloads come only from your tutor or this site. Never from a pop-up, an error page, a forum, or a stranger's video. The "paste this command to fix it" scam is the single most common way real people get hacked, and it works by getting you to do the pasting.
  2. Never type a real password into anything Claude-adjacent. Not into the terminal, not into a chat. Signing in happens in your browser, on the real website, and nowhere else.
  3. Remember: anything Claude reads can steer what Claude does. So the question "where should this project live?" is really "will Claude be reading things I didn't write and don't trust?", and if yes, it belongs somewhere Claude has walls, or no hands at all. The table below does this sorting for you.
  4. Protect the accounts themselves. Your GitHub and Claude accounts are part of your safety. Two cheap habits defeat almost all real-world account attacks: a password you use nowhere else, and two-step sign-in (where the site also checks your phone) turned on.

Which project needs which safety?

We decide by asking a question: how risky is this thing?

Note: Risk lives in what a project reads, runs and touches, not in the person doing it. A stranger's webpage is exactly as dangerous for a professional as for a beginner.

What you're doing How risky? Why it's ranked that way Best home
The course itself LOW Tiny exercises with nothing at stake, and mistakes are part of the point. CAGE
A personal website, portfolio, or browser game, published via GitHub LOW Public by design, nothing private involved. Worst case is a broken page, and Git undoes it. CAGE
Reading things from the internet for you: "summarise this article", "compare these two products" MEDIUM Web pages you didn't write can carry hidden instructions that try to steer the AI. CLAUDE WEBSITE: in a plain chat, the worst a steered AI can do is give you a bad answer
Home hardware: a Raspberry Pi, smart-home tinkering LOW Stakes are naturally small: mistakes affect a hobby gadget, not your machine. OUTSIDE THE CAGE: it needs your home network, which the cage deliberately blocks

Worth noticing: the two columns answer different questions, and they don't always move together. The Raspberry Pi row is low risk yet lives outside the cage, because the home is decided by what a project needs to reach, and a Pi needs your home network, which the cage blocks. Risk tells you how careful the setup has to be; what the project needs to touch tells you which setup that is.

The day you get your first API key.

Sooner or later a project will want an API key: a long secret code that lets your project use a paid service, where every use costs money on your card. That day, four things:

Unsure where a project belongs?

Research first: the WhatsApp group while you're with us, someone experienced, or a plain chat with your AI. Ask before you start, not after.