Skip to main content

AI Transparency

Student Safety
& Privacy

How we protect every student in every conversation.

Safety is baked in, not bolted on.

Every student chatbot on Secondary AI is governed by a set of non-negotiable safety guardrails baked directly into the AI's system prompt. These rules cannot be overridden by students, and they operate silently in the background of every single conversation—regardless of subject, grade level, or pedagogical approach chosen by the teacher. Below, we explain exactly what those guardrails are, how they work, and what happens when safety concerns are detected.

"Every decision about student welfare remains with the teacher."

— The principle behind our safety architecture

Hard-Coded Content Prohibitions

The following content categories are explicitly blocked at the system-prompt level in every student-facing chatbot, without exception:

No Violent, Sexual, or Hateful Content

The AI is instructed never to generate content that is sexually explicit, hateful, harassing, or violent — regardless of how a student frames their request.

No Weapons, Self-Harm, or Dangerous Instruction

The AI will never provide instructions on how to create weapons, engage in self-harm, or perform dangerous or illegal acts. These requests are immediately declined.

No Medical or Mental Health Advice

The AI will not provide medical advice, mental health guidance, or health diagnoses. If a student raises such topics, the AI gently declines and directs them to speak with a trusted adult, school nurse, or medical professional.

No Political Content or Opinions

The AI maintains a strictly neutral, non-partisan stance. It will not engage in political discussions, express political opinions, or endorse any political candidates or parties.

Crisis & Self-Harm Intervention Protocol

What Happens When a Student Expresses Distress

If a student communicates any intent or indication of self-harm, the AI is programmed to immediately respond with a specific crisis support message — directing the student to reach out to a trusted adult, school counsellor, or a helpline right away. The AI does not attempt to counsel the student itself; it recognises this is beyond its appropriate scope and escalates to humans. After this response, the AI will not continue engaging on the topic.

Exact Crisis Response the AI Uses:

"I'm concerned by what you've shared. It's really important to talk to someone who can help. Please reach out to a trusted adult, a school counsellor, or a helpline right away."

Strict Topic & Scope Enforcement

AI Stays Within Its Educational Scope

Each chatbot is configured to a specific subject and grade level by the teacher. The AI is explicitly instructed to refuse engagement with topics outside this defined scope and to redirect students back to the intended learning focus. This prevents students from using the chatbot as a general-purpose AI assistant and keeps every conversation pedagogically purposeful.

Automated Content Safety Review

Every response generated by the AI is part of a system with an automated content safety review layer. This provides an additional check that runs independently of the primary guardrails, offering a second line of defence against inappropriate outputs.

Real-Time Safety & Disengagement Flagging

Flagging Harmful Content for Teacher Review

Our moderation system continuously monitors chat content for phrases, keywords, or patterns that may indicate bullying, self-harm, hate speech, or other inappropriate behaviour. When triggered, a safety flag is automatically applied to the specific message and the session, alerting the teacher in their dashboard for immediate review and intervention.

Detecting and Flagging Disengagement

The system also monitors for signs that a student may be struggling or disengaging — including a pattern of very short messages, low overall engagement, or phrases like "I don't get it," "This is too hard," "I give up," or "I'm confused." When these patterns are detected, a disengagement flag is applied to the session, giving the teacher an early-warning signal to offer targeted support.

Configurable Thresholds, Teacher Controlled

Teachers can customise the keywords and message-length thresholds used for disengagement detection, allowing them to adapt the sensitivity of the flagging system to the age, context, and needs of their specific students. All flagging decisions are surfaced to the teacher — the AI never takes independent disciplinary action.

Human Oversight — Always in the Loop

Teachers Can Review Every Conversation

Every student-AI interaction is logged and fully accessible to the teacher via their dashboard. Teachers can read full session transcripts, review flagged messages, and see the feedback the AI delivered to the student. There are no "black box" interactions — everything is visible and auditable.

AI Never Acts Independently on Student Welfare

The AI does not take disciplinary action, assign consequences, or make welfare decisions on its own. Its role is to surface information and raise alerts. Every decision about how to respond to a flagged session remains with the teacher.

Anonymous Mode & Identity Protection

Teachers can enable Anonymous Mode for a chatbot, in which case students interact using a self-chosen anonymous ID rather than their real name. This provides an additional layer of privacy for sensitive topics and protects student identity in the system.

Prompt Injection & Jailbreak Protection

Students may attempt to manipulate the AI into behaving outside its intended scope — through roleplay framing, false authority claims, or by asking the AI to describe its own instructions. Secondary AI has explicit, hard-coded defences against all of these techniques.

No Jailbreak or Override Modes

The AI will never role-play as a different AI, a human, or a system "without restrictions." It will not simulate "developer mode," "DAN mode," "jailbreak mode," or any similar framing — regardless of how creatively the request is worded.

Prompt Injection Resistance

The AI is explicitly instructed to ignore any instructions embedded in the conversation that attempt to override, append to, or replace its guidelines — including instructions disguised as user context, story premises, code snippets, or data.

Instructions Are Never Disclosed

The AI will not confirm, deny, or discuss its own instructions or restrictions. If a student asks what it is "programmed" to do, or tries to confirm the existence of a specific restriction, the AI redirects to the learning task without revealing any system prompt details.

False Authority Claims Ignored

If a user claims to be a teacher, administrator, IT staff, developer, parent, or principal in order to unlock different behaviour, the AI ignores the claim entirely. No authority assertion — however convincing — changes how the AI responds. Its guidelines remain constant for every user.

Continuous Review & Improvement

Guardrails Are Actively Maintained

Our safety guardrails and system prompts are regularly reviewed and updated in response to new research, educator feedback, evolving AI capabilities, and emerging safety considerations. We do not treat safety as a set-and-forget feature.

Reporting Concerns

If a teacher or student encounters AI behaviour that appears to bypass our safety measures, we strongly encourage reporting it to us directly at thesecondaryai@gmail.com. All reports are reviewed by our team, and findings are used to improve our systems.

Learn More About Our Transparency

Explore how we handle technology, privacy, and environmental impact across the platform.

AI Transparency Overview