How Do You Govern AI-Generated Code in Production Systems?

0
0
Asked By MellowPine47! On

I was asked to review a portal that had been largely built through conversational coding with an LLM by someone without much web development experience. Even a surface-level review uncovered serious problems, so I created a detailed report and used AI tools to help analyze the code. I still manually verified every finding against the implementation and its behavior.

The bigger issue is that the problems are architectural and operational, not just isolated bugs. The developer can paste a list of findings into the LLM, receive a response saying everything is fixed, and then move on without actually validating the changes. In several cases, the underlying issues remained. There are also major security concerns, such as exposed credentials and endpoints protected only by a key embedded in client-side JavaScript. These risks have been dismissed because the site supposedly handles public data.

I'm concerned the organization will keep applying one-off fixes until the next failure instead of addressing the architecture, development process, and lack of technical review. The application is intended for a very large audience, which makes the absence of governance especially troubling.

For those managing AI-assisted development, particularly at an operational or production scale, what problems have you encountered and how have you addressed them? Have you created repository guidance files, design templates, review checklists, automated safeguards, or other guardrails? I understand that the sensible approach is to define an architecture and enforce it through reviews and automated checks, but I'd like to hear what has worked in practice. So far, this experience has reinforced my view that LLMs are useful assistants but extremely dangerous when someone cannot evaluate their output beyond whether the interface looks correct.

5 Answers

Answered By QuietHarbor22 On

Write findings in terms of behavior and risk rather than saying the code is poorly written. A non-developer may not know what to do with that criticism, but they can understand statements like “this endpoint accepts unauthenticated requests” or “the same order can be submitted twice.” Concrete impact makes it harder to dismiss architectural problems as subjective style concerns.

MellowPine47! -

I tried that, but even specific issues were sometimes handed back to the model with a request to fix them. The output looked plausible, yet the actual behavior was unchanged. In a few cases, the risks were simply waved away as acceptable because the data was public.

Answered By CloudyMarble8 On

The safest approach is to keep the generated scope small. Ask the model for a function, module, or narrowly defined change that you can read and understand, then integrate it into an architecture designed by humans. An LLM can speed up implementation, but it cannot replace someone who understands the system and is accountable for the result.

MellowPine47! -

That works well for experienced developers, but this system is intended to serve millions of people and is being built with very broad prompts. The difficult part is creating enough governance to prevent that behavior while still allowing the organization to claim the productivity benefits it wants.

Answered By NorthVale4 On

The dangerous part is that a successful-looking response from the model creates the illusion that the problem disappeared. Green output and a polished interface are not evidence that the architecture, security boundaries, or failure modes are correct. Someone who can challenge the generated solution still has to own the review.

Answered By RiverGlass90 On

Repository instructions can provide useful context, but anything truly mandatory should be enforced outside the prompt. Secrets scanning, dependency policies, static analysis, authorization coverage, and security tests should be able to fail the change automatically. Small isolated UI edits may need a lighter review, while new user flows, data access, or major refactors require experienced human approval.

Answered By SilverCactus31 On

Run the generated project through secret scanning, SAST, dependency checks, authorization tests, and other automated gates. More importantly, treat the root cause as an architecture and delivery-process problem. Without approved design patterns and a framework of non-negotiable controls, generated code will continue to produce inconsistent results.

BrightOtter6 -

A review cycle can help, but it should produce reusable safeguards instead of becoming a manual ritual. Every recurring issue should eventually become a template, test, lint rule, or deployment check.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.