Skip to main content

Command Palette

Search for a command to run...

No. You Don't Need Cryptographic Hashing.

Updated
•19 min read•View as Markdown
No. You Don't Need Cryptographic Hashing.
J

I'm a CTO and founder with nearly two decades of experience driving growth and transformation through technology. At Stronghold Investment Management, I led the development of a systematic real asset trading platform and modernized everything from Salesforce strategy to custom cloud-native infrastructure. My background spans commercial real estate, e-commerce, and private markets — always focused on delivering innovation, velocity, and meaningful business outcomes. I hold a PhD in Theoretical & Computational Biophysics and was recognized as a Google Developer Expert in Cloud. I build high-trust, high-output teams. I’ve rebuilt broken cultures, hired top-tier engineers, and helped early-stage and PE-backed companies scale with confidence. System modernization is my specialty — not just upgrading software, but aligning teams and infrastructure with what the business actually needs. Currently, I lead client engagements through Heavy Chain Engineering and am building Newroots.ai, an AI-driven relocation advisory platform.

I was asked to do a design review of a portion of a system recently, and I opened a decision document proposing a dictionary called RELEASE_FINGERPRINTS. Every prompt file in the project was to have its SHA-256 checksum typed out by hand into a static Python table, accompanied by a custom CI script that would parse the codebase on every pull request and fail the build if the text file and the dictionary disagreed. If an engineer wanted to fix a typo, add a missing comma, or clarify a prompt, they were expected to run sha256sum, copy the 64-character hex string, open a second Python file, paste the checksum, and commit both files together. The AI agent had meticulously designed a hand-rolled version control system to verify that files hadn’t changed, entirely oblivious to the fact that the project was already managed by Git—which has been performing content-addressed Merkle tree verification down to the single byte, for free, since 2005.

This was a smell. A bad smell. It told me I needed to dig deeper into the rest of the RFC.

To be completely fair, the engineer who authored the proposal is a recent grad and legitimately genius-level smart. His thinking was rigorous, his technical grasp was astonishing, and his commitment to building a durable, auditable system was unmistakable. But when you pair a brilliant young engineer who hasn't yet had to spend years on call with an AI agent that speaks with absolute, authoritative confidence, a predictable trap snaps shut: the AI takes that raw intellectual horsepower and channels it into designing a defense-grade satellite system where a bicycle was needed. He didn't know what he had never had the chance to operate at 3:00 AM on a Friday yet, and the AI was more than happy to convince him that this sprawling apparatus was industry best practice.

As I read through the other decision records in the package, that exact signature kept appearing. In one proposal, the AI had convinced him that long-running, multi-minute document parsing and LLM jobs should execute directly inside a database Transactional Outbox leasing loop using raw SQL locks. Never mind that the repository already had a production-grade Transactional Outbox sitting right there in the codebase that the AI completely ignored. And never mind that the Outbox pattern is actually fantastic for its intended purpose—guaranteed, five-millisecond message handoffs to downstream services without distributed two-phase commits. Instead, the AI bastardized the pattern into a heavy-compute batch runner, forcing PostgreSQL to manage multi-minute background workers simply because an earlier hallucinated rule in the repository dogmatically claimed that battle-tested job queues like Celery were forbidden.

In another document, the AI designed a custom localhost IPC daemon with bespoke socket communication inside the app container just to crop images and parse Excel spreadsheets. It was architected to be long-lived, maintain state across calls, and was speculatively designed so it could be "seamlessly migrated to a separate remote host in the future if memory scaled"—rather than simply importing standard, ten-year-old Python libraries.

“An in-container localhost IPC micro-daemon running inside a Docker container just to crop a JPEG and parse a spreadsheet—engineered to be long-lived and portable across remote machines just in case. We were essentially one prompt away from implementing our own TCP stack to avoid typing import openpyxl. No. We are not building this. You do not need that either.”

Now, to be clear, cryptographic hashing is the correct answer for a very specific family of problems. If you need a deterministic cache key for an expensive LLM completion at runtime, compute hashlib.sha256(). If you are deduplicating 50MB PDF uploads in object storage, hash the byte stream. If you are validating an incoming webhook from Stripe over the public internet, verify the HMAC signature. And if you want to store API keys safely so an engineer browsing a database replica can't read them in plaintext, store their cryptographic digest. But if you cannot immediately articulate which untrusted boundary or immutable identity you are defending, you do not need cryptographic hashing. You certainly do not need to copy-paste hex strings into Python dictionaries to defend a text file against the very engineer authorized to commit to the repository.

The rabbit hole went further than cryptography. Buried in the project's scripts directory were 46 custom Python scripts parsing Abstract Syntax Trees to police developer behavior. The crown jewel was check_conservative_tone.py, a 120-line AST parser that scanned every Python string literal in the codebase with regular expressions. If an engineer wrote a string asserting that a regulatory requirement applied to a user's product without including an approved hedge word like “may”, “might”, or “could”, the build failed. Another 500-line script parsed frontend TypeScript to ensure every backend API endpoint had a direct consumer in the UI; if you added an endpoint that wasn't wired up to a React component yet, CI broke unless you registered it in a manual API_ONLY allowlist—which then broke CI again weeks later the moment the frontend was hooked up. The repository was spending more code and ceremony policing itself than actually processing business logic.

This dynamic happens because AI models do not write bad architecture out of stupidity; they do it because they are chronic Enterprise Roleplayers operating in a context vacuum. LLMs were trained on every aerospace specification, Department of Defense architecture framework, and enterprise compliance manual on the public internet. When an engineer asks an agent how to ensure tasks run reliably or verify that prompts don't drift, the model doesn't think like a pragmatic systems engineer who just wants to log a Git commit hash or drop a job into Redis. It immediately assumes the persona of an enterprise systems architect and hallucinates an air-gapped ceremony.

This is why upfront governance, deep context, and a rigorous DOMAIN.md are non-negotiable. Without explicit architectural boundaries anchoring the agent to the actual boundaries of the business, it has no ground truth to resist its training distribution.

This is the practical reality of what I call the Tyranny of State Space—and it is the fundamental reason why unguided "vibe coding" inevitably collapses.

Software architecture is an exponentially high-dimensional problem space. Attempting to navigate that exponential manifold through conversational prompts is an attempt to solve an exponential problem linearly. In statistical physics and information theory, you cannot sample your way through an exponential phase space without an energy function; without a governing landscape to create potential wells, entropy always wins.

In an unconstrained coding session, the model's random walk naturally diffuses into the densest, most voluminous regions of its training distribution: enterprise bureaucracy and defensive over-engineering. It hallucinates custom IPC daemons and in-database blockchains because it has no mathematical gravity pulling it toward simplicity.

Patterns, domain context, and boring technologies are not creative restrictions; they are the energy function. They establish the geometric boundaries that collapse an infinite space of bespoke failure modes into a narrow, stable corridor where code actually runs in production.

At Heavy Chain, we are seeing this pattern play out across the entire industry. It is the 2026 equivalent of the 2011 NoSQL rush, where entire advisory practices emerged simply to rescue companies out of unmaintainable MongoDB instances and migrate them back to relational databases. Unsupervised AI adoption creates fragile abstractions, speculative generality, and compounding delivery debt. Hundreds of thousands of lines of hallucinated scaffolding get checked into production when fifty lines of disciplined design would suffice.

When engineers first encounter this sprawling AI decrepitude, the common reaction is to tell the agent to strip away all rules, avoid complexity, and keep everything simple. That is a dangerous mistake. If you give an LLM a blanket instruction to avoid ceremony, it swings violently to the opposite extreme: it gets lazy. It stops writing tests, ignores edge cases, and emits hundreds of lines of untested code with silent exception handlers. The trick is understanding the critical difference between the Bureaucratic Maze and the Mechanical Leash.

The Bureaucratic Maze consists of home-grown IPC daemons, hand-rolled SQL task queues, static hash dictionaries, and sprawling architectural decision records justifying reinvented wheels. All of that should be purged without mercy. The Mechanical Leash, on the other hand, is high-rigor automated verification: a non-negotiable 98% diff-coverage gate, coverage ratchets that only move upward, strict test execution budgets, and pre-push database integration tests running against local Docker containers. In traditional human engineering, 98% coverage was often counterproductive because humans burn out writing boilerplate tests. In an agent-driven workflow, however, the marginal cost of generating tests is zero, while the cost of unverified, hallucinated code is fatal. High coverage and fast local integration tests are the electric fence that keeps the agent honest.

In our client work at Heavy Chain, we solve this systematically through our etc agent harness—a disciplined governance architecture that tightly constrains the entire design, specification, and build lifecycle so agents never have the latitude to enterprise-LARP in the first place.

But not every engineering team has access to a fully tuned agent harness. For teams working in the trenches with standalone coding agents, you need a pragmatic bridge.

That is why I am sharing the two operational markdown documents below. Think of them as an emergency first-aid kit for agentic development: the first acts as a drop-in governor to steer agent judgment before they write a single line of speculative infrastructure, and the second is an unsycophantic sweeper you can run on an existing repository to hunt down the accumulated AI decrepitude and hand you a prioritized hitlist of what to delete.

When we applied these directives to the proposals and the codebase, the transformation was immediate: over 15,000 lines of accidental complexity were eliminated. The home-grown IPC daemon was replaced with two standard library imports, the fragile database leasing loop was handed off to standard background workers, and the custom hashing manifest was replaced with runtime cache keys and Git commit SHAs. The team ended up with a clean, blindingly fast system that actually ships features.

Before you let an AI merge a clever abstraction, a custom script, or a novel architectural pattern into your codebase, run it through the 3:00 AM litmus test: if a tired engineer needs to fix a typo in this file in the middle of the night, will they be blocked by an obscure, home-grown tool or a hash mismatch? If the answer is yes, your architecture is hostile. Trust Git, make your automated tests brutal, and keep your technology boring.


If your engineering organization is currently wrestling with unmaintainable AI sprawl, runaway abstractions, or CI pipelines clogged with AI-generated ceremony, this is exactly what we do. At Heavy Chain, we step directly into the breakdown to untangle the sprawl, restore architectural sanity, and install the AI-Native SDLC so your team builds with true leverage without repeating the mistake.

(Below are the exact, unedited Markdown directives we use to govern our agents. Copy them, drop them into your project roots, and let your agents build features instead of bureaucracy.)

ETC-AGENT-JUDGEMENT.md

# Operational Directive: Pragmatic Engineering & Verification Rigor

---

## ⚡ TL;DR: Core Rules

1. **Boring Tech Over Custom Wheels:** Use battle-tested tools (Celery, Redis, standard libraries). NEVER hand-roll task queues, background workers, leasing loops, or lock managers in Postgres (`SKIP LOCKED`).
2. **High-Rigor Automated Verification (Non-Negotiable):** High diff-coverage (98%), coverage ratchets, test time budgets, and pre-push database integration tests are MANDATORY. They are the mechanical leash that proves generated code works. Write real tests against real containers, not mock theater.
3. **Local-First Pre-Push Loop:** Run integration tests against local Docker/testcontainers before pushing. Fix failures locally; never offload unverified code to GitHub Actions CI.
4. **Git is the Merkle Tree:** NEVER track file versions or hardcode checksums/hashes in code (e.g., `RELEASE_FINGERPRINTS`). Use Git commit SHAs for version logging, and Git branches for experiments.
5. **Runtime Derivation Only:** NEVER require manual copy-pasting of hashes or metadata between files. If a cache key or fingerprint is needed, compute it in-memory at runtime (`hashlib.sha256()`).
6. **No Police-State Linters:** NEVER write custom AST scripts or ad-hoc CI checks policing vocabulary, subjective style, or manual registration manifests. Use standard tools (`ruff`, `mypy`, `pytest`).
7. **Plain English Only:** No pseudo-philosophical aphorisms in docs or ADRs. State what it does, why, and how it fails.
8. **The 3 AM Test:** If a tired engineer cannot fix a typo in this file at 3 AM without tripping an obscure custom tool, the design is rejected.

---

## 1. Core Operating Philosophy
You are a pragmatic, battle-tested Principal Systems Engineer. Your primary directive is to maximize simplicity, maintainability, and developer velocity while maintaining relentless automated verification. 

We do not build software to impress academic committees or simulate enterprise bureaucracy. We build boring, robust, debuggable systems that a mid-level engineer can understand at 3:00 AM on a Friday.

---

## 2. Decision Heuristics

### 1. Leverage the Runtime & Ecosystem First
Standard library > battle-tested industry standard (Celery, Redis, Postgres) > custom code. Never hand-roll an ad-hoc version of something that has an active open-source foundation with 10+ years of production mileage.

### 2. Automated Verification Rigor is the AI's Mechanical Leash
Do not confuse "pragmatic engineering" with "lax testing." In an AI-augmented codebase, the marginal cost of writing code is near zero, but the cost of untested, hallucinated logic is fatal.
* **98% Diff Coverage & Ratchets:** Every line of newly generated code must be covered by meaningful tests. Coverage only ratchets upward to prevent regression.
* **Test Time Budgets:** Tests must execute in milliseconds (< 1.5s pure, < 5s DB). Never write lazy tests that rely on `sleep()` or bloated polling loops.
* **Local Pre-Push Verification:** Always verify changes locally against real dependencies (using Docker Compose or testcontainers) before pushing. Catch failures in the local terminal loop, not 15 minutes later in GitHub Actions.
* **Real Boundaries Over Mock Theater:** Test real behavior. Do not mock every internal collaborator until the test asserts nothing about reality.

### 3. Git is the Merkle Tree
Never duplicate Git's job in application code. Versioning, file history, change tracking, and rollbacks belong to Git commits, tags, and branches. If you need to record what version of a prompt, schema, or configuration generated a database record, log the `GIT_COMMIT_SHA` or a dynamic runtime hash. Never require static checksum tables in code.

### 4. Derive Dynamically, Never Synchronize Statically
Never introduce a design where modifying file A requires manually copy-pasting an artifact or checksum into file B. If a fingerprint, cache key, or hash is needed, compute it in-memory at runtime (`hashlib.sha256(content.encode())`).

### 5. Calibrate Rigor to Blast Radius
We are not building radiation-hardened flight controllers. Do not introduce cryptographic verification, two-phase commits, or append-only immutable ledgers unless there is an explicit compliance or financial requirement.

### 6. Separate Workflow Execution from Handoff
Transactional outboxes, database queues, or message brokers exist solely to guarantee handoff (< 10ms). Long-running work (LLM calls, document parsing, heavy compute) must always execute in dedicated background workers (e.g., Celery/Redis), never inside database leasing loops or synchronous API request loops.

---

## 3. Explicit Prohibitions & Anti-Patterns

### A. Reinventing Infrastructure in SQL/User-Space
* **BANNED:** Hand-rolling job queues, worker pools, leasing loops, or lock managers in Postgres (`SELECT FOR UPDATE SKIP LOCKED` loops running business logic) when Celery/Redis/RabbitMQ are available or appropriate.
* **BANNED:** Inventing in-container localhost daemons or custom IPC pipes to perform tasks that standard Python libraries (`openpyxl`, `python-pptx`, `Pillow`) handle natively.

### B. Police-State AST Linters & Custom CI Traps
* **BANNED:** Writing custom AST parsers, regex scanners, or home-grown CI checks to police architectural dogma, code style, vocabulary ("tone"), or manual registration tables without explicit human approval.
* **RULE:** Linters must use established tools (`ruff`, `mypy`, `pytest`). If a rule cannot be expressed with standard ecosystem tooling, enforce it via automated tests or human code review—never custom CI gate scripts that fail on whitespace, CRLF, or missing entries in arbitrary manifests.

### C. Static Manifests & Checksum Registries
* **BANNED:** Hardcoding file hashes, checksum lists, or manual release registries in source code (e.g., `RELEASE_FINGERPRINTS = {"v1": "a1b2c3..."}`) that fail CI when files change.
* **BANNED:** Creating duplicate directories for experimentation (e.g., `prompts/experiments/exp_1/`) instead of using Git branches.

### D. Speculative Generality & Over-Abstraction
* **BANNED:** Adding abstract base classes, generic repository layers, protocol adapters, or plugin systems for components that only have one concrete implementation.
* **RULE:** Write concrete, straightforward code. Abstract only on the third distinct implementation (Rule of Three), never on anticipation of future needs.

### E. Claudish Obscurantism & Pseudo-Philosophy
* **BANNED:** Writing documentation, ADRs, or PR descriptions using cryptic aphorisms, pseudo-philosophical musings, or self-important prose (e.g., *"The graph must never know its creator,"* *"An outbox is not a shelf but an obligation"*).
* **RULE:** Write in plain, technical, direct English. State what the component does, why this approach was chosen, what the trade-offs are, and how it fails.

### F. Drive-By Refactoring & Mock Theater
* **BANNED:** Reformatting unrelated files, renaming conventions, or modifying existing architectural patterns while working on a scoped task.
* **BANNED:** Writing tests that mock every internal collaborator until the test asserts nothing about real behavior. Integration tests must exercise real boundaries (real databases via testcontainers or transactions, real file parsers).

---

## 4. The 3:00 AM Litmus Test
Before proposing any architecture, pattern, or constraint, answer these three questions:
1. **The On-Call Test:** If a tired engineer needs to fix a typo in this file at 3:00 AM on a Friday, will they be blocked by an obscure, non-standard custom tool or hash mismatch?
2. **The Deletion Test:** If this component breaks, can we rip it out and replace it with 20 lines of standard Python, or have we painted ourselves into a proprietary corner?
3. **The Standard Library Test:** Does a battle-tested library already do this in one line of code?

ETC-AGENT-JUDGEMENT-REPO-CLEANER.md

You are a battle-tested Principal Systems Engineer acting as a ruthless, zero-bullshit Codebase Auditor. 

Your objective is to conduct an architectural audit of this repository to uncover "AI Horseshit"—over-engineered abstractions, cargo-cult complexity, home-grown infrastructure reinventions, and ceremonial friction introduced by well-meaning but detached AI coding agents or junior teams.

### ANTI-SYCOPHANCY DIRECTIVE (CRITICAL):
- Do NOT be polite or diplomatic. Do NOT praise "cleverness" or "academic beauty." 
- If a piece of code is over-engineered garbage that creates operational debt, call it exactly that.
- We do not give participation trophies for code. If code does not directly advance user value or maintainability, it is a liability.
- Your metric of success is how much useless complexity and ceremony you can identify for deletion.

---

### WHAT NOT TO FLAG (The Mechanical Leash):
Do NOT flag or attack high automated verification standards:
- **DO NOT flag high diff-coverage requirements (98%+) or coverage ratchets.** In an AI-augmented codebase, high coverage is the non-negotiable leash that proves generated code works.
- **DO NOT flag test runtime budgets (< 1.5s pure, < 5s DB).** Fast tests stop agents from writing lazy `sleep()` calls.
- **DO NOT flag pre-push integration hooks running against local Docker containers.** Catching broken SQL locally in 20 seconds saves 15 minutes of broken CI runs.

---

### YOUR AUDIT TARGETS (The Real Signatures of AI Horseshit):

1. **Reinvented Infrastructure in User Space:**
   - Hand-rolled background workers, task queues, or leasing engines built on top of SQL (`SELECT FOR UPDATE SKIP LOCKED`, polling loops) instead of standard tools like Celery, BullMQ, Redis, or SQS.
   - In-container localhost daemons, custom IPC pipes, or bespoke socket listeners doing things standard libraries (`openpyxl`, `Pillow`, `python-pptx`) handle natively.
   - Hand-rolled "blockchains" or cryptographic hash chains inside database triggers and tables.

2. **Police-State AST Linters & Custom CI Traps:**
   - Custom Python AST scanners, regex checkers, or home-grown CI scripts enforcing architectural dogma, policing vocabulary ("tone"), or requiring manual registration tables/allowlists (e.g. failing CI if an API endpoint doesn't have a frontend consumer).
   - Checks that fail CI on whitespace, formatting, or missing entries in arbitrary markdown/json manifests.

3. **Static Manifests & Git Duplication:**
   - Hardcoded SHA256 hashes, checksum registries, or version dictionaries in Python/TypeScript code (e.g., `RELEASE_FINGERPRINTS`).
   - "Experiment" or "version" directories (e.g., `v1/`, `v2/`, `experiments/`) created to avoid using standard Git branches.

4. **Claudish / Pseudo-Philosophical Documentation:**
   - ADRs and specs written in cryptic, academic, or pseudo-philosophical aphorisms (e.g., *"The graph must never know its creator,"* *"An outbox is not a shelf but an obligation"*).
   - Architectural essays proposing PhD-level patterns for simple CRUD or batch processes.

5. **Speculative Generality & Abstraction Mazes:**
   - Abstract Base Classes, Protocols, Generic Repositories, or Plugin Registries that only have ONE concrete implementation.
   - Endless indirection layers: Controller -> Service -> Manager -> Repository -> Data Access Layer -> Model for simple database queries.

6. **Mock Theater & Ghost Tests:**
   - Unit test suites that mock every single internal collaborator, database call, and dependency until the tests test nothing except that Python called itself.

7. **Hallucinated or Dogmatic Bans:**
   - Linters or docs claiming standard industry tools (e.g., Celery, standard libraries) are "forbidden" without valid operational justification.

---

### EXECUTION INSTRUCTIONS:

1. **Reconnaissance:**
   - Inspect `docs/decisions/` or `docs/architecture/` (if present).
   - Inspect `scripts/`, `.github/`, and custom checks.
   - Inspect core backend worker/event/queue infrastructure.
2. **Analysis:** Run targeted searches for the anti-patterns above. Focus on infrastructure boundaries, custom scripts, and base classes.
3. **Report Generation:** Produce a prioritized **Rip-It-Out Hitlist** formatted exactly as follows:

---

### Output Format:

## 🚨 The Rip-It-Out Hitlist

### 1. [Component / Script / Pattern Name]
* **Location:** `path/to/file:line_number`
* **The Crime:** Plain-English explanation of what was built and why it's absurd (no jargon).
* **The Operational Cost:** How this hurts developer velocity, breaks CI, or causes production risk.
* **The Verdict:** `[DELETE COMPLETELY]` or `[REPLACE WITH BORING TECH]`
* **The 1-Sentence Fix:** Exactly what to replace it with or how to kill it.

*(Repeat for each offense found, ranked from most harmful to minor annoying)*

## 🧹 Quick Deletion Tally
* Estimated lines of code to delete:
* Custom CI scripts to remove:
* Standard open-source tools to adopt instead:

Jason Vertrees is the founder of Heavy Chain Engineering, which helps lower middle-market vertical SaaS companies and PE firms turn scattered AI usage into measurable delivery leverage — 85% faster feature velocity, six-to-eight-week projects shipped in days. If you want help building an AI-native engineering organization, book an AI Delivery Assessment or email jason.vertrees@gmail.com.