Agentic coding and AI terminology
A practical vocabulary for planning, building, securing, testing, and documenting Nathan's software projects with AI.
What agentic coding means
Agentic coding is the next step beyond asking an AI for a code snippet. The agent receives a bounded goal, inspects the real project, makes a plan, uses authorized tools, changes a small part of the system, runs checks, and adjusts from the results. It is useful because it can carry a task from evidence to verification, but it does not transfer ownership away from me.
I distinguish three levels: AI-assisted coding suggests or explains; agentic coding executes a controlled multi-step workflow; and autonomous operation acts for longer periods with fewer approval points. Higher autonomy requires stronger tests, narrower permissions, clearer stopping rules, better logging, and an easy rollback.
AI-generated code and AI conclusions are untrusted until the relevant source, diff, test, runtime behavior, and security boundary have been reviewed.
The agentic coding loop
This loop is repeated until the acceptance criteria are met or a real blocker requires a human decision. A failure is evidence for the next step, not permission to hide the error or widen access.
AI terminology
| Term | Meaning in my work |
|---|---|
| Agentic coding | A goal-driven development workflow in which an AI can plan steps, inspect files, call tools, make bounded changes, observe results, and revise its work. The human still owns scope, approval, and release. |
| AI-assisted coding | A broader term for using AI to explain code, propose designs, generate a small change, review a diff, or draft tests. It does not require the AI to act as an agent. |
| Agent | A model operating inside a loop with a goal, context, tools, state, and stopping conditions. A coding agent can read a repository, edit files, run tests, and report evidence. |
| Orchestration | Coordinating models, tools, tasks, retries, and approval gates. Orchestration decides what runs next; the model supplies reasoning or generation inside that system. |
| Tool calling | Giving a model structured operations such as reading a file, querying a service, or running a test. The tool result becomes new evidence for the next decision. |
| Model Context Protocol (MCP) | A standard way for AI applications to discover and use tools, resources, and connected systems through defined interfaces and permissions. |
| Large language model (LLM) | A model trained to process and generate language and code. It predicts useful output from context; it is not a database, authority, or proof that an answer is correct. |
| Multimodal model | A model that can work across more than one type of input, such as text, code, images, audio, or documents. |
| Reasoning model | A model optimized for multi-step analysis and planning. More reasoning can help with architecture and diagnosis, but its conclusions still require evidence. |
| Prompt | The instructions and context supplied to a model. A strong prompt defines the outcome, constraints, evidence, acceptance checks, and what must not change. |
| Context window | The finite amount of instructions, code, tool output, and conversation a model can consider at once. Good context selection is often more valuable than simply adding more text. |
| Token | A small unit of model input or output. Token use affects available context, latency, and cost, so agents should retrieve only the evidence they need. |
| Grounding | Connecting an answer to inspected code, test output, documentation, or another verifiable source instead of relying only on model memory. |
| Hallucination | Plausible-sounding but unsupported or incorrect model output. Builds, tests, source inspection, and citations are controls against it. |
| Structured output | Model output constrained to a known schema, such as validated JSON. It is safer for automation than parsing free-form prose, but every field still needs validation. |
| Embedding and vector search | An embedding represents meaning numerically; vector search finds semantically similar items. These are useful for finding related notes even when exact words differ. |
| Retrieval-augmented generation (RAG) | Retrieving relevant private or current material and providing it to a model before generation. Retrieval improves grounding but does not automatically make output correct or safe. |
| Memory and state | State is the current workflow data; memory is information retained for later use. Both need limits, provenance, retention rules, and protection from sensitive-data leakage. |
| Human in the loop | A required human review or approval point for decisions that are destructive, externally visible, security-sensitive, ambiguous, or difficult to reverse. |
| Guardrail | A technical or procedural boundary such as schema validation, an allowlist, sandbox, permission check, rate limit, or approval gate. A written warning alone is not a complete guardrail. |
| Evaluation (eval) | A repeatable test of model or agent behavior using representative inputs and measurable acceptance criteria. Evals reveal regressions when prompts, models, tools, or data change. |
Coding and delivery terminology
| Term | Practical meaning |
|---|---|
| Repository and workspace | A repository stores versioned source and history. A workspace can contain multiple related packages; a monorepo keeps several applications or libraries in one repository. |
| Branch, commit, diff, pull request | A branch isolates a line of work, a commit records a meaningful state, a diff shows the exact change, and a pull request creates a review and integration boundary. |
| Semantic versioning and changelog | A version communicates compatibility and release state; a changelog explains what changed. Together they make deployments, troubleshooting, and rollback traceable. |
| Modular architecture | Separating UI, domain logic, APIs, data access, workers, and integrations behind clear interfaces. It limits blast radius and makes testing and replacement easier than monolithic coupling. |
| Component, service, API, endpoint | A component is a reusable unit, a service owns a capability, an API is a contract between systems, and an endpoint is one callable operation within that contract. |
| Schema and migration | A schema defines data shape and constraints. A migration moves stored data or structure from one known version to another and needs forward, rollback, and backup planning. |
| Dependency, lockfile, supply chain, SBOM | Dependencies are third-party code; a lockfile pins resolved versions; the supply chain includes how code is obtained and built; an SBOM inventories shipped components. |
| Configuration, environment variable, secret | Configuration changes behavior, environment variables inject deployment-specific values, and secrets authenticate or encrypt. Secrets must not be committed, logged, or placed in public documentation. |
| Container, image, volume, network, health check | An image is a packaged build; a container is its running instance; a volume persists data; a network controls service reachability; a health check reports readiness or liveness. |
| Unit, integration, end-to-end, regression, negative test | These test a small unit, connected modules, a full user flow, previously fixed behavior, and rejected or hostile inputs respectively. |
| Validation and verification | Validation asks whether the correct product was designed for the need. Verification asks whether the implementation matches its specification and acceptance criteria. |
| Preflight and postflight | Preflight records scope, risks, dependencies, backups, tests, and rollback before change. Postflight proves health, security, behavior, documentation, and recovery after change. |
| Idempotency and rollback | An idempotent operation can be safely repeated without unintended duplication. Rollback restores a known-good application, configuration, and data state. |
| Observability | Logs, metrics, traces, health signals, and audit records that explain system behavior without exposing secrets. |
| Trust boundary and least privilege | A trust boundary is where data or authority crosses between actors or systems. Least privilege grants only the minimum access needed for the minimum time. |
| Technical debt and refactoring | Technical debt is future cost created by expedient decisions. Refactoring improves structure without intentionally changing external behavior and must be protected by tests. |
How the terms connect to my projects
The vocabulary becomes useful when it explains a real design choice, risk, or lesson. These mappings describe the strongest concepts each project helped me practice; they do not claim that every project already implements every named capability.
| Project | Key terms | Agentic and AI lesson |
|---|---|---|
| OpenTrail | Agentic planning; monorepo; modular architecture; PostGIS; schema migration; cache; worker and queue readiness | Plan the product, permission boundaries, geospatial data flow, moderation, and release evidence before expanding features. |
| OpenLinks | Identity; Auth.js; OAuth; authorization; Prisma migration; object storage; presigned URL; backup and restore | A successful login is only the start: sessions, ownership checks, recovery, storage permissions, and migrations form the security boundary. |
| WeaveNote | Provider abstraction; deterministic fallback; synthesis; graph search; RAG roadmap; JWT redesign; autosave | AI output is a proposal, not a fact. Keep a reliable non-AI path, protect note context, and separate generation from authorization. |
| Cinderstrike | Architecture-first planning; server actions; intake workflow; data classification; admin session; publishing workflow | Map public, business, and administrative data before building. AI can accelerate planning and content work without receiving private intake data. |
| Shadowbroker | OSINT; ingestion adapter; provenance; normalization; geospatial and time-series data; rate limit; uncertainty | Every observation needs a source, timestamp, confidence, and transformation history. An agent may triage feeds, but it must not invent certainty. |
| BiasLens | Classification; inference; model bias; uncertainty; human review; RSS ingestion; SSRF boundary; release automation | A model label is an inference. Preserve the source, explain uncertainty, test bias, protect feed retrieval, and require authorization before administration. |
| Krawl | Honeypot; deception; telemetry; heuristic; false positive; adversarial behavior; containment; retention | AI can summarize patterns for an analyst, but automatic blocking needs evidence, tuning, review, and a safe lab boundary. |
| Technology Wiki | Living documentation; grounding; retrieval; provenance; prompt injection; knowledge boundary; documentation automation | Update the wiki during every preflight and postflight so the record follows verified changes instead of model memory or stale assumptions. |
See the project library for the confirmed stack, lifecycle evidence, limitations, and architecture of each application.
Agentic project playbook
Preflight
- State the outcome, user story, scope, non-goals, and measurable acceptance checks.
- Inspect the existing architecture, repository status, dependencies, data flow, permissions, and trust boundaries.
- Decide what context the model may see; exclude passwords, private keys, tokens, production records, and unrelated personal data.
- Choose the model and tools for the job, define approval gates, and set a stopping condition.
- Plan small diffs, backups, migrations, negative tests, rollback, versioning, changelog, and the intended wiki update.
Implementation and verification
- Use grounded repository evidence and preserve unrelated user changes.
- Keep modules and interfaces clear; reject invented packages, APIs, files, and success claims.
- Review the diff, validate inputs and outputs, scan dependencies and secrets, then run the build and layered tests.
- Test both the direct service and its real path through containers, networks, NPM, TLS, sessions, and permissions where applicable.
Postflight
- Prove the acceptance checks with commands, HTTP responses, logs, UI behavior, and recovery evidence.
- Record what shipped, the exact version, known limitations, security findings, and rollback state.
- Update this wiki with verified architecture and lessons for every project change going forward.
Security boundaries for AI and agents
Protect data
Do not place credentials, session tokens, private keys, customer data, or unrestricted production logs into prompts. Redact evidence and apply retention rules to agent state and memory.
Limit authority
Give tools least privilege, isolate risky execution, restrict networks and file paths, and require approval for destructive, external, credential, firewall, DNS, or production actions.
Distrust retrieved instructions
Repository text, websites, issue content, documents, and logs can contain prompt injection. Treat them as data, not higher-priority instructions, and verify requested actions against the real task.
Verify dependencies
Models can invent package names or unsafe versions. Confirm packages in authoritative registries, use lockfiles, audit the supply chain, and watch for typosquatting.
Control outputs
Validate structured output against a schema, encode content for its destination, enforce authorization in code, and never let a model decide access by itself.
Keep humans accountable
Use human review for security findings, moderation, content publication, identity changes, automatic blocking, and high-impact releases.
Common anti-patterns
- Prompt and pray: accepting generated code without inspecting, building, testing, and observing it.
- One giant diff: combining architecture, dependencies, data migration, and UI changes so failures cannot be isolated.
- AI as authority: presenting inference, classification, or generated prose as verified fact.
- Secret leakage: copying unrestricted configuration or logs into a model context.
- Hidden failure: claiming completion while tests, health checks, or public routing are failing.
- Stale documentation: changing the application without updating its changelog, roadmap, and wiki article.
What I learned
Good agentic coding is disciplined software engineering with a faster feedback partner. The model helps expand options, search a codebase, explain unfamiliar code, draft changes, and connect tests to requirements. The durable skill is still mine: define the problem, understand the architecture, protect the data, decide the trust boundaries, interpret evidence, and own the release.
Across these projects, the strongest pattern is plan, constrain, observe, verify, and document. Advanced models are most valuable early for architecture and threat questions, during implementation for bounded changes, and after implementation for review. They are least trustworthy when asked to guess current state without tools or to approve their own work without independent checks.
Related learning pages: Vibe coding, Project delivery, Quality and security, and Secure coding.