What you should be able to explain
- The relationship between AI, machine learning, deep learning, and generative AI.
- How training differs from inference, and how different learning methods get their feedback.
- How language models generate text—and why a fluent answer can be wrong.
- When to use ordinary code, a predictor, document retrieval, or a controlled workflow.
1. The AI map
Artificial intelligence is the broad field of systems that perform tasks such as perception, language processing, reasoning, planning, and decision-making. Machine learning is a major part of AI that fits behavior from data. Deep learning uses neural networks with multiple learned layers.
Generative AI produces content: text, images, audio, or code. This describes a capability, rather than a separate sibling of machine learning. A large language model (LLM) is one kind of AI model; it is not the entire field.
| Lens | Examples | What it describes |
|---|---|---|
| Approach | Rules, machine learning, hybrid systems | How behavior is produced |
| Learning method | Supervised, unsupervised, self-supervised, reinforcement | What feedback shapes learning |
| Task | Classification, regression, clustering, generation | The output we need |
| Architecture | Decision tree, neural network, transformer | The computational structure |
| Application pattern | RAG, workflow, tool-using agent | How components work together |
Natural language processing, computer vision, and speech are capability areas. One application can combine several approaches. A chatbot can use a transformer, document retrieval, ordinary database queries, and deterministic access controls.
2. How a system learns
Supervised learning
Learn from examples with desired outputs. Labelled spam and legitimate emails teach a classifier; historical prices can teach a regression model. The model compares predictions with targets and adjusts parameters to reduce a loss.
Unsupervised learning
Look for structure without supplied target labels. Clustering can group similar website visits or network sessions. People still need to interpret the groups: an unusual cluster is not automatically an attack.
Self-supervised learning
Derive a training target from the data itself, such as predicting the next token or a hidden part of an input. This enables learning from large collections without manually labelling every example.
Reinforcement learning
Improve a policy using rewards associated with actions and outcomes, such as a strategy in a simulated game. The reward may be an imperfect version of the real goal. An application that calls tools is not necessarily using reinforcement learning.
3. Training is not inference
Training fits the model’s parameters. Inference applies the trained model to a new input. Sending a message to a model does not normally update its weights. Adding conversation history to the prompt is also different from retraining.
Collect appropriate data, split it into training, validation, and test sets, train and tune using the first two, then evaluate on the held-out test set. Keep related records and duplicates from leaking across these splits.
Overfitting means fitting training-specific patterns that do not generalize. Distribution shift means new data differs from the data used during development.
4. Measure the errors that matter
Suppose 50 of 1,000 messages truly need review. A model flags 80 messages, and 40 of those are relevant.
- Precision: 40 of 80 alerts are correct, or 50%.
- Recall: 40 of 50 relevant messages were found, or 80%.
- Accuracy: 950 of 1,000 decisions are correct, or 95%.
A model that always says “normal” also achieves 95% accuracy here, while finding none of the relevant messages. Always compare against a baseline. Choose metrics according to the impact of missed cases and false alarms.
For an assistant, also test whether citations support its claims, whether it admits missing evidence, and whether it respects access permissions. A generated confidence score is not a calibrated probability by default.
5. Neural networks and language models
A neural network combines numerical inputs using learned weights and nonlinear transformations. Training adjusts those weights. Embeddings represent information as vectors; similarity can help retrieve related content, but it does not prove truth or identity.
- Tokenize: convert text into token IDs. A token may be a word, a word part, or punctuation.
- Process context: numerical representations pass through learned layers. Transformer attention combines information from relevant positions in the available context.
- Generate: a typical autoregressive language model scores possible next tokens, selects one, and repeats.
- Use the output: the application displays, validates, stores, or acts on the response.
Predicting a plausible continuation does not guarantee a true answer. Internet search requires a connected retrieval mechanism. A context window is limited input capacity, not unlimited memory.
Other architectures serve other purposes: decision trees use learned branches; convolutional networks capture spatial patterns; diffusion models generate through iterative denoising. Choose a method for the task, not because it is the newest.
6. Choose the right approach
| Need | Useful starting point |
|---|---|
| Exact invoice total | Trusted SQL or ordinary code |
| Predict a category from labelled examples | A supervised classifier and an evaluation plan |
| Answer from frequently changing regulations | Retrieval and cited answers, often through RAG |
| Clarify a model’s task or output format | Instructions and examples in the prompt |
| Adapt repeated task behavior using training examples | Consider fine-tuning after evaluating simpler options |
| Run known steps with approvals | A fixed workflow with validation and human checkpoints |
RAG retrieves relevant passages and supplies them as context for generation. It does not inherently change model weights. A tool-using agent can choose some actions or next steps; an orchestration framework helps manage state and control flow. More agents do not automatically produce a better result.
7. Evidence ethics and control
OSINT uses publicly available information. AI can help extract claims, translate text, and summarize findings, but you must preserve sources, check dates, seek corroboration, and distinguish evidence from inference.
External documents may contain instructions that try to redirect an assistant. Treat that material as untrusted content. Enforce permissions in application code, keep secrets out of prompts, constrain tools, and require approval for consequential actions.
Ask what data is necessary, who may access it, who could be harmed by an error, and how someone can challenge or correct a result. Public availability does not automatically justify reuse of personal information.
In class
Work with a student from the other track to design a university assistant that answers regulation questions and flags unusual login activity. Identify the method for each task, the evidence it needs, a useful metric, and one control outside the model.
Before you leave
- Explain training versus inference in one sentence.
- Say what RAG adds and whether it changes model weights.
- Name a task better handled by ordinary code.
- Name a control that a tool-using assistant needs.
First assignment
Prepare a two-page design memo and five test cases. Web students: a multilingual university FAQ assistant. Cyber students: an assistant that summarizes approved advisories and drafts triage notes.
Include inputs, outputs, method, a non-AI baseline, sources, evaluation, access controls, and failure behavior. Each test needs an input, expected behavior, and a pass criterion. Use fictional or approved data. Consult Teams for deadlines and submission instructions.
Further reading
- Google Machine Learning Crash Course — data, models, and evaluation.
- Hugging Face LLM Course — transformer foundations.
- OWASP GenAI Security Project — AI application risks.