How should I structure a full-stack AI chat application using Next.js, TypeScript, PostgreSQL, and an LLM API? #209527
🏷️ Discussion TypeQuestion BodyWhat I'm trying to buildI am building a full-stack AI chat application using Next.js, React, TypeScript, PostgreSQL, and an LLM API. The application will have:
My current planned stack is:
My current architectureI am thinking about structuring the application like this: For PDF/document processing, I am thinking about: My main programming questionI am not sure how I should properly organize the code. I am considering a structure like: I would like to know if this is a good structure for a production application. Chat request flowWhen a user sends a message, I want something similar to: I am thinking about keeping the API route relatively small and moving the actual logic into separate functions/services. For example: POST /api/chatcould internally call: Is this a good programming pattern? LLM provider architectureI also want to avoid tightly coupling the entire application to one AI provider. I was thinking about creating an interface such as: interface AIProvider {
generateResponse();
streamResponse();
createEmbedding();
}Then having implementations such as: This would allow me to change AI providers later without rewriting the entire application. Is this a good approach, or is it unnecessary abstraction for an MVP? Database architectureI also want to understand the correct way to access the database. Should the frontend directly access the database? Or should it always go through server-side code? I want to make sure that users can only access their own:
and can never access another user's data. AI MemoryI want the AI to have long-term memory. For example, if the user tells the AI: I may want the application to save this as a memory. I am considering: Then later: Is this a good architecture for AI memory? RAGFor document questions, I want to use RAG. For example: Then: I would also like the answer to reference the source page of the PDF. For example: What is the recommended way to store page numbers and metadata so that I can generate reliable citations? StreamingI want the AI response to appear progressively in the UI instead of waiting for the complete response. For example: then: then: What is the recommended way to implement streaming with Next.js and TypeScript? Should I use:
for an AI chat application? Authentication and securitySecurity is very important because users will have private conversations and documents. I want the architecture to ensure: is impossible. I am considering authentication through Supabase Auth / Auth.js / Clerk. For database security, I am also considering PostgreSQL Row Level Security. Is this enough, or are there other important security controls I should implement? I also understand that LLM API keys should never be exposed in the browser. The flow should be: and not: Guidelines
|
Replies: 3 comments
|
Your architecture is generally correct. For a first production MVP, I would not start with Next.js + FastAPI + Redis + multiple microservices. That adds complexity before you need it. Use a modular monolith: 1. Project structureYour structure is good. I would organize it approximately like this: The important principle is:
2.
|
|
Your proposed architecture is very solid and aligns well with current best practices for full-stack AI apps. Next.js with Server Actions is a great fit for handling the API layer, and for streaming LLM responses, you should definitely look into the Vercel AI SDK if you haven't already — it handles the streaming UI hooks seamlessly. For the database layer, using PostgreSQL for your relational data (Users, Conversations, Messages, metadata) is exactly the right call. Regarding the RAG, Document Chunks, and AI Memory components: starting with If you want to keep the operational simplicity of not running a separate search cluster but want that advanced retrieval, you might want to look at Infino. It's an open-source embedded retrieval engine that runs in-process (like SQLite) and has Node.js bindings. It handles vector kNN, full-text (BM25), and hybrid search out of the box. You could use Postgres for your user/app state, and run Infino embedded in your Next.js backend specifically for the document chunks and AI memory retrieval. For file storage, the standard pattern is to drop the raw PDFs into an S3-compatible object store, parse them in your backend (Next.js API route or a background worker), and then index the resulting chunks into your retrieval engine. Disclosure: I work on Infino. |
Root Cause Analysis & Immediate Preflight Check WorkaroundHi @wscha231! Your report and API traces provide a textbook breakdown of a backend state-machine desync in GitHub Actions. 🔍 Why HTTP 409 Occurs on Cancel/Force-CancelThis happens due to an internal divergence between the Workflow Run aggregate state and the Job Execution coordinator:
Because user-facing APIs enforce validation against the job scheduler, there are no public API endpoints or UI toggles that can forcibly write a terminal 🛠️ 1. Official Route for State ReconciliationTo reconcile the run state without deleting logs, artifacts, or execution history:
💡 2. Immediate Client-Side Workaround for CI Preflight ChecksWhile waiting for backend reconciliation, you can adjust your preflight script so these historical ghost runs don't block your deployment pipeline: Option A: Filter by Age Threshold (Recommended)Add a timestamp filter in your preflight check to ignore runs older than a realistic execution timeout (e.g., created > 24 hours ago): # Example jq filter ignoring orphaned runs older than 24 hours
gh api "repos/wscha231/r1000-quant-engine/actions/runs?status=in_progress" | \
jq --arg cutoff "$(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ)" \
'[.workflow_runs[] | select(.created_at > $cutoff)]' |
Your architecture is generally correct. For a first production MVP, I would not start with Next.js + FastAPI + Redis + multiple microservices. That adds complexity before you need it.
Use a modular monolith:
1. Project structure
Your structure is good. I would organize it approximately like this: