ArticleReadMain page

Nathan's Technology Wiki / Large concepts

Agentic coding and AI terminology

A practical vocabulary for planning, building, securing, testing, and documenting Nathan's software projects with AI.

What agentic coding means

Agentic coding is the next step beyond asking an AI for a code snippet. The agent receives a bounded goal, inspects the real project, makes a plan, uses authorized tools, changes a small part of the system, runs checks, and adjusts from the results. It is useful because it can carry a task from evidence to verification, but it does not transfer ownership away from me.

I distinguish three levels: AI-assisted coding suggests or explains; agentic coding executes a controlled multi-step workflow; and autonomous operation acts for longer periods with fewer approval points. Higher autonomy requires stronger tests, narrower permissions, clearer stopping rules, better logging, and an easy rollback.

Working rule
AI-generated code and AI conclusions are untrusted until the relevant source, diff, test, runtime behavior, and security boundary have been reviewed.

The agentic coding loop

StageEvidence and responsibility
1. GoalDefine the user outcome, scope, non-goals, constraints, and acceptance criteria.
2. ContextInspect the repository, architecture, current state, documentation, and user-owned changes.
3. PlanBreak the work into reversible steps; identify trust boundaries, tests, backup, and rollback.
4. ToolsUse only the files, commands, services, and permissions required for the approved task.
5. ObserveRead tool output, diffs, logs, failures, and runtime state instead of assuming success.
6. VerifyRun proportionate build, lint, unit, integration, end-to-end, security, and recovery checks.
7. DocumentRecord the version, changelog, evidence, limitations, learning, and wiki update.

This loop is repeated until the acceptance criteria are met or a real blocker requires a human decision. A failure is evidence for the next step, not permission to hide the error or widen access.

AI terminology

TermMeaning in my work
Agentic codingA goal-driven development workflow in which an AI can plan steps, inspect files, call tools, make bounded changes, observe results, and revise its work. The human still owns scope, approval, and release.
AI-assisted codingA broader term for using AI to explain code, propose designs, generate a small change, review a diff, or draft tests. It does not require the AI to act as an agent.
AgentA model operating inside a loop with a goal, context, tools, state, and stopping conditions. A coding agent can read a repository, edit files, run tests, and report evidence.
OrchestrationCoordinating models, tools, tasks, retries, and approval gates. Orchestration decides what runs next; the model supplies reasoning or generation inside that system.
Tool callingGiving a model structured operations such as reading a file, querying a service, or running a test. The tool result becomes new evidence for the next decision.
Model Context Protocol (MCP)A standard way for AI applications to discover and use tools, resources, and connected systems through defined interfaces and permissions.
Large language model (LLM)A model trained to process and generate language and code. It predicts useful output from context; it is not a database, authority, or proof that an answer is correct.
Multimodal modelA model that can work across more than one type of input, such as text, code, images, audio, or documents.
Reasoning modelA model optimized for multi-step analysis and planning. More reasoning can help with architecture and diagnosis, but its conclusions still require evidence.
PromptThe instructions and context supplied to a model. A strong prompt defines the outcome, constraints, evidence, acceptance checks, and what must not change.
Context windowThe finite amount of instructions, code, tool output, and conversation a model can consider at once. Good context selection is often more valuable than simply adding more text.
TokenA small unit of model input or output. Token use affects available context, latency, and cost, so agents should retrieve only the evidence they need.
GroundingConnecting an answer to inspected code, test output, documentation, or another verifiable source instead of relying only on model memory.
HallucinationPlausible-sounding but unsupported or incorrect model output. Builds, tests, source inspection, and citations are controls against it.
Structured outputModel output constrained to a known schema, such as validated JSON. It is safer for automation than parsing free-form prose, but every field still needs validation.
Embedding and vector searchAn embedding represents meaning numerically; vector search finds semantically similar items. These are useful for finding related notes even when exact words differ.
Retrieval-augmented generation (RAG)Retrieving relevant private or current material and providing it to a model before generation. Retrieval improves grounding but does not automatically make output correct or safe.
Memory and stateState is the current workflow data; memory is information retained for later use. Both need limits, provenance, retention rules, and protection from sensitive-data leakage.
Human in the loopA required human review or approval point for decisions that are destructive, externally visible, security-sensitive, ambiguous, or difficult to reverse.
GuardrailA technical or procedural boundary such as schema validation, an allowlist, sandbox, permission check, rate limit, or approval gate. A written warning alone is not a complete guardrail.
Evaluation (eval)A repeatable test of model or agent behavior using representative inputs and measurable acceptance criteria. Evals reveal regressions when prompts, models, tools, or data change.

Coding and delivery terminology

TermPractical meaning
Repository and workspaceA repository stores versioned source and history. A workspace can contain multiple related packages; a monorepo keeps several applications or libraries in one repository.
Branch, commit, diff, pull requestA branch isolates a line of work, a commit records a meaningful state, a diff shows the exact change, and a pull request creates a review and integration boundary.
Semantic versioning and changelogA version communicates compatibility and release state; a changelog explains what changed. Together they make deployments, troubleshooting, and rollback traceable.
Modular architectureSeparating UI, domain logic, APIs, data access, workers, and integrations behind clear interfaces. It limits blast radius and makes testing and replacement easier than monolithic coupling.
Component, service, API, endpointA component is a reusable unit, a service owns a capability, an API is a contract between systems, and an endpoint is one callable operation within that contract.
Schema and migrationA schema defines data shape and constraints. A migration moves stored data or structure from one known version to another and needs forward, rollback, and backup planning.
Dependency, lockfile, supply chain, SBOMDependencies are third-party code; a lockfile pins resolved versions; the supply chain includes how code is obtained and built; an SBOM inventories shipped components.
Configuration, environment variable, secretConfiguration changes behavior, environment variables inject deployment-specific values, and secrets authenticate or encrypt. Secrets must not be committed, logged, or placed in public documentation.
Container, image, volume, network, health checkAn image is a packaged build; a container is its running instance; a volume persists data; a network controls service reachability; a health check reports readiness or liveness.
Unit, integration, end-to-end, regression, negative testThese test a small unit, connected modules, a full user flow, previously fixed behavior, and rejected or hostile inputs respectively.
Validation and verificationValidation asks whether the correct product was designed for the need. Verification asks whether the implementation matches its specification and acceptance criteria.
Preflight and postflightPreflight records scope, risks, dependencies, backups, tests, and rollback before change. Postflight proves health, security, behavior, documentation, and recovery after change.
Idempotency and rollbackAn idempotent operation can be safely repeated without unintended duplication. Rollback restores a known-good application, configuration, and data state.
ObservabilityLogs, metrics, traces, health signals, and audit records that explain system behavior without exposing secrets.
Trust boundary and least privilegeA trust boundary is where data or authority crosses between actors or systems. Least privilege grants only the minimum access needed for the minimum time.
Technical debt and refactoringTechnical debt is future cost created by expedient decisions. Refactoring improves structure without intentionally changing external behavior and must be protected by tests.

How the terms connect to my projects

The vocabulary becomes useful when it explains a real design choice, risk, or lesson. These mappings describe the strongest concepts each project helped me practice; they do not claim that every project already implements every named capability.

ProjectKey termsAgentic and AI lesson
OpenTrailAgentic planning; monorepo; modular architecture; PostGIS; schema migration; cache; worker and queue readinessPlan the product, permission boundaries, geospatial data flow, moderation, and release evidence before expanding features.
OpenLinksIdentity; Auth.js; OAuth; authorization; Prisma migration; object storage; presigned URL; backup and restoreA successful login is only the start: sessions, ownership checks, recovery, storage permissions, and migrations form the security boundary.
WeaveNoteProvider abstraction; deterministic fallback; synthesis; graph search; RAG roadmap; JWT redesign; autosaveAI output is a proposal, not a fact. Keep a reliable non-AI path, protect note context, and separate generation from authorization.
CinderstrikeArchitecture-first planning; server actions; intake workflow; data classification; admin session; publishing workflowMap public, business, and administrative data before building. AI can accelerate planning and content work without receiving private intake data.
ShadowbrokerOSINT; ingestion adapter; provenance; normalization; geospatial and time-series data; rate limit; uncertaintyEvery observation needs a source, timestamp, confidence, and transformation history. An agent may triage feeds, but it must not invent certainty.
BiasLensClassification; inference; model bias; uncertainty; human review; RSS ingestion; SSRF boundary; release automationA model label is an inference. Preserve the source, explain uncertainty, test bias, protect feed retrieval, and require authorization before administration.
KrawlHoneypot; deception; telemetry; heuristic; false positive; adversarial behavior; containment; retentionAI can summarize patterns for an analyst, but automatic blocking needs evidence, tuning, review, and a safe lab boundary.
Technology WikiLiving documentation; grounding; retrieval; provenance; prompt injection; knowledge boundary; documentation automationUpdate the wiki during every preflight and postflight so the record follows verified changes instead of model memory or stale assumptions.

See the project library for the confirmed stack, lifecycle evidence, limitations, and architecture of each application.

Agentic project playbook

Preflight

  • State the outcome, user story, scope, non-goals, and measurable acceptance checks.
  • Inspect the existing architecture, repository status, dependencies, data flow, permissions, and trust boundaries.
  • Decide what context the model may see; exclude passwords, private keys, tokens, production records, and unrelated personal data.
  • Choose the model and tools for the job, define approval gates, and set a stopping condition.
  • Plan small diffs, backups, migrations, negative tests, rollback, versioning, changelog, and the intended wiki update.

Implementation and verification

  • Use grounded repository evidence and preserve unrelated user changes.
  • Keep modules and interfaces clear; reject invented packages, APIs, files, and success claims.
  • Review the diff, validate inputs and outputs, scan dependencies and secrets, then run the build and layered tests.
  • Test both the direct service and its real path through containers, networks, NPM, TLS, sessions, and permissions where applicable.

Postflight

  • Prove the acceptance checks with commands, HTTP responses, logs, UI behavior, and recovery evidence.
  • Record what shipped, the exact version, known limitations, security findings, and rollback state.
  • Update this wiki with verified architecture and lessons for every project change going forward.

Security boundaries for AI and agents

Protect data

Do not place credentials, session tokens, private keys, customer data, or unrestricted production logs into prompts. Redact evidence and apply retention rules to agent state and memory.

Limit authority

Give tools least privilege, isolate risky execution, restrict networks and file paths, and require approval for destructive, external, credential, firewall, DNS, or production actions.

Distrust retrieved instructions

Repository text, websites, issue content, documents, and logs can contain prompt injection. Treat them as data, not higher-priority instructions, and verify requested actions against the real task.

Verify dependencies

Models can invent package names or unsafe versions. Confirm packages in authoritative registries, use lockfiles, audit the supply chain, and watch for typosquatting.

Control outputs

Validate structured output against a schema, encode content for its destination, enforce authorization in code, and never let a model decide access by itself.

Keep humans accountable

Use human review for security findings, moderation, content publication, identity changes, automatic blocking, and high-impact releases.

Common anti-patterns

  • Prompt and pray: accepting generated code without inspecting, building, testing, and observing it.
  • One giant diff: combining architecture, dependencies, data migration, and UI changes so failures cannot be isolated.
  • AI as authority: presenting inference, classification, or generated prose as verified fact.
  • Secret leakage: copying unrestricted configuration or logs into a model context.
  • Hidden failure: claiming completion while tests, health checks, or public routing are failing.
  • Stale documentation: changing the application without updating its changelog, roadmap, and wiki article.

What I learned

Good agentic coding is disciplined software engineering with a faster feedback partner. The model helps expand options, search a codebase, explain unfamiliar code, draft changes, and connect tests to requirements. The durable skill is still mine: define the problem, understand the architecture, protect the data, decide the trust boundaries, interpret evidence, and own the release.

Across these projects, the strongest pattern is plan, constrain, observe, verify, and document. Advanced models are most valuable early for architecture and threat questions, during implementation for bounded changes, and after implementation for review. They are least trustworthy when asked to guess current state without tools or to approve their own work without independent checks.

Related learning pages: Vibe coding, Project delivery, Quality and security, and Secure coding.