Moodly is a wellness app. A user logs how they feel right now (typing a mood, taking a photo, or recording a short voice note), and the app uses that to suggest a real activity or "experience" they can book: a guided meditation, a group walk, a workshop, a community event. Hosts create these experiences, users discover and book them, and everyone can join communities built around shared interests.
This repository is the business backend: accounts, bookings, experiences, communities, notifications, and the logic that turns a detected mood into a recommendation. It does not run the AI models itself. A separate FastAPI service (a different repository) does the actual emotion detection and text embedding; this backend calls out to it and reacts to the result.
If you're joining this project, the fastest way to get oriented is: read this page top to bottom once, then open the README inside whichever module you're about to touch, every folder under src/ has one with its exact API and internal wiring.
- Architecture
- Project Structure
- Module Index
- Backend Base Structure (how a module is wired)
- API Versioning
- Features
- Engineering Challenges Handled
- Event-Driven Architecture (RabbitMQ)
- Background Jobs (BullMQ / Bull)
- Sample Request Flow: Creating a Booking
- Sample Request Flow: Mood Log Detection and Recommendation
- Database
- Scale: Current Capacity and Where Overflow Goes
- Auth
- FastAPI Inference Service (external repo)
- Environment Variables
- Project Setup
- Deployment
- Observability
- Known Gaps / Next Planned Work
Style: a modular monolith split into two deployable processes, connected by an event bus. This is not a microservices architecture: both processes share one codebase, one set of entities, and the same Postgres/Mongo databases. It's closer to "one application, two runtimes":
- API process (
src/main.ts, port3002): everything that answers an HTTP request. Must stay fast and responsive. - Worker process (
src/worker/main.ts, port3001): everything slow, retryable, or dependent on an external AI call. Runs as a separate OS process (npm run start:worker), consuming events from RabbitMQ.
The two are decoupled through RabbitMQ (durable, at-least-once event delivery) rather than direct function calls or HTTP, so a slow AI inference call in the worker can never block someone's booking request in the API. Real-time updates back to the client go over Socket.IO. Delayed and retryable work that stays inside the API process (like sending an email) goes through BullMQ instead of RabbitMQ, see Background Jobs for why there are two different queue mechanisms.
flowchart TB
Client["Client app"] -->|"REST + cookie or JWT"| API["NestJS API process<br/>port 3002"]
Client -->|"Upload photo or voice"| API
API -->|"HTTP call"| FastAPI["FastAPI inference service<br/>port 8000, separate repo"]
API -->|"Publish event"| Queue["RabbitMQ<br/>5 exchange and queue pairs"]
Queue -->|"Event consumed"| Worker["NestJS Worker process<br/>port 3001"]
Worker -->|"Publish follow-up event"| Queue
Worker -->|"Socket.IO push"| Client
API --> Postgres["PostgreSQL<br/>via TypeORM"]
Worker --> Postgres
API --> Mongo["MongoDB<br/>via Mongoose"]
Worker --> Mongo
API -->|"Enqueue job"| Redis["Redis<br/>BullMQ queues"]
API --> S3["AWS S3<br/>images and avatars"]
API --> Email["Email via Nodemailer"]
Stack at a glance: NestJS 11 (TypeScript) on both processes, PostgreSQL via TypeORM for relational domain data, MongoDB via Mongoose for onboarding documents and vector embeddings, RabbitMQ for cross-process events, Redis for BullMQ job queues, Socket.IO for real-time push, JWT + Google OAuth for auth, AWS S3 for file storage, and an external FastAPI service for the actual AI models.
ai-moodler-backend/
├── src/
│ ├── main.ts # API process bootstrap (port 3002)
│ ├── app.module.ts # composition root, imports every feature module
│ ├── app.controller.ts / app.service.ts
│ │
│ ├── auth/ # JWT (access+refresh) + Google OAuth
│ ├── users/ # user CRUD, plus profile/ (self-service, avatar, password)
│ ├── onboarding/ # multi-step onboarding, Mongo-backed, user + host variants
│ ├── experience/ # bookable "experience" domain, host/public/user controllers
│ ├── booking/ # booking lifecycle, race-safe creation, host dashboards
│ ├── attendance/ # QR / join-code check-in, one-to-one with booking
│ ├── mood-log/ # mood logging (text/photo/voice), starts the mood event chain
│ ├── embedding/ # text embedding generation + Mongo vector storage
│ ├── recommendation/ # emotion + embedding based matching, optional LLM rerank
│ ├── feedback/ # post-experience ratings, cron-driven reminder queue
│ ├── notification/ # in-app, email, and Socket.IO notifications
│ ├── community/ # groups, posts, reactions, comments
│ ├── insights/ # aggregated user analytics
│ ├── diagram/ # GET /v1/diagram, live Mermaid module graph
│ │
│ ├── common/ # cross-cutting: S3, Gemini, FastAPI client, ResultDto, roles guard
│ ├── logger/ # Winston config, AsyncLocalStorage request id
│ ├── database/ # TypeORM + Mongoose module, migrations, data-source
│ │
│ ├── infra/ # technical plumbing, not domain logic (see src/infra/README.md)
│ │ ├── redis/ # RedisService: ioredis wrapper, cache + lock primitives
│ │ ├── rmq/ # RmqModule: RabbitMQ producer client factory
│ │ ├── config/ # RMQ_DOMAINS: exchange/queue/routing-key map
│ │ └── bull-board/ # mounts the Bull Board admin UI at /admin/queues
│ │
│ └── worker/ # SEPARATE process entrypoint (main.ts, port 3001)
│ ├── embedding.worker.ts
│ ├── recommendation.worker.ts
│ ├── onboarding.worker.ts
│ └── experience.worker.ts
│
├── certs/rds-ca.pem # RDS Postgres SSL cert (production only)
├── test/jest-e2e.json # e2e harness config (no specs committed yet)
├── init.sql # manual schema reference (pgcrypto + vector extensions)
├── AGENTS.md, CLAUDE.md # instructions for AI coding agents working in this repo
├── nest-cli.json, tsconfig*.json, eslint.config*.mjs
└── package.json
Every module below has its own README.md with the full endpoint reference: method, route, guards, request/response shape. This root document covers cross-cutting architecture only.
| Module | Path | README |
|---|---|---|
| Auth | src/auth |
src/auth/README.md |
| Users | src/users |
src/users/README.md |
| Profile | src/users/profile |
src/users/profile/README.md |
| Onboarding | src/onboarding |
src/onboarding/README.md |
| Experience | src/experience |
src/experience/README.md |
| Booking | src/booking |
src/booking/README.md |
| Attendance | src/attendance |
src/attendance/README.md |
| Mood Log | src/mood-log |
src/mood-log/README.md |
| Embedding | src/embedding |
src/embedding/README.md |
| Recommendation | src/recommendation |
src/recommendation/README.md |
| Feedback | src/feedback |
src/feedback/README.md |
| Notification | src/notification |
src/notification/README.md |
| Community | src/community |
src/community/README.md |
| Insights | src/insights |
src/insights/README.md |
| Diagram | src/diagram |
src/diagram/README.md |
| Common | src/common |
src/common/README.md |
| Database | src/database |
src/database/README.md |
| Worker process | src/worker |
src/worker/README.md |
| Infra (overview) | src/infra |
src/infra/README.md |
| Redis | src/infra/redis |
src/infra/redis/README.md |
| RabbitMQ client | src/infra/rmq |
src/infra/rmq/README.md |
| Config | src/infra/config |
src/infra/config/README.md |
| Bull Board | src/infra/bull-board |
src/infra/bull-board/README.md |
Every feature module follows the same NestJS layering, for example booking:
Controller (HTTP layer: guards, DTO validation, HTTP status mapping)
|
Service (orchestrator, e.g. BookingService)
|
Specialized services (single responsibility: *-creation, *-validation,
*-query, *-mapper, *-side-effects, *-error-handler, *-stats ...)
|
TypeORM Repository / Mongoose Model -> PostgreSQL / MongoDB
Conventions used across the codebase:
- Guards stack: most protected routes use
@UseGuards(JwtBearerGuard, JwtCookieGuard, RolesGuard).JwtCookieGuardalone already reads thejwthttpOnly cookie or falls back to anAuthorization: Bearerheader, soJwtBearerGuardis redundant on routes that also applyJwtCookieGuard, kept for explicitness in most controllers. - Response envelope: almost every handler returns
ResultDto.ok(data, message, statusCode)or throws anHttpExceptionbuilt fromResultDto.fail(...), giving a consistent{ success, statusCode, data, message }/{ success: false, statusCode, reason, errorType }shape across the whole API (src/common/dto/result.dto.ts,src/common/constants/error-code-map.ts). - Validation: global
ValidationPipe({ transform: true, whitelist: true })(src/main.ts); every DTO usesclass-validatordecorators. - Fire-and-forget side effects: write-path services (e.g.
BookingCreationService) commit the DB transaction first, then dispatch notifications and attendance creation as detached promises (.catch()-guarded, not awaited), so a slow email or notification path never adds latency to the HTTP response. - Split controllers per audience: several domains (
experience,booking,onboarding) expose separatehost/*,user/*, andpublic/*controllers instead of one controller with internal role branching, keeping route-level guards and Swagger tags unambiguous.
Every HTTP route in this API is served under /v1/... via NestJS's built-in URI versioning (app.enableVersioning({ type: VersioningType.URI, defaultVersion: '1' }) in src/main.ts). POST /auth/login from an older version of this README (or any client hardcoded before this change) is now POST /v1/auth/login, and so on for every route documented in this file and every module README.
This exists to protect anyone building against this API (a mobile app, a future webhook consumer, an internal admin tool) from breaking changes: a future incompatible change to, say, the booking response shape can ship as a new /v2/user/bookings route on a specific controller (@Controller({ path: 'user/bookings', version: '2' })) while every existing /v1/... integration keeps working unmodified, no big-bang cutover required.
Two routes are deliberately excluded (version: VERSION_NEUTRAL) rather than versioned, following the same convention most production APIs use:
GET /(AppController), a bare root route intended for load balancer health checks and uptime monitors that shouldn't need to know the current API version to hit it.- Swagger UI (
/api-docs) and Bull Board (/admin/queues) aren't affected either way, they're mounted directly rather than through a versioned Nest controller, so they were never prefixed and still aren't.
Everything else, including the dev-only /v1/diagram introspection route, gets the default /v1/ prefix automatically, nothing had to change per-controller. Socket.IO gateways are untouched by this: enableVersioning only applies to HTTP routes, not WebSocket connections.
Regression-tested in test/versioning.e2e-spec.ts (an isolated check against throwaway controllers, not the full AppModule, since that needs a reachable Postgres/Mongo/Redis/RabbitMQ that a fast unit-style e2e test shouldn't depend on): confirms a route with no explicit version is served at /v1/, is not reachable without the prefix, and that a controller opting into version: '2' is served at /v2/ instead.
| Feature | Description |
|---|---|
| Email/Password Auth | bcrypt-hashed (cost 12) signup and login with access/refresh token rotation |
| Google OAuth 2.0 | Passport-based social login, find-or-create by email |
| Role-based Access | user, host, admin roles enforced via RolesGuard + @Roles() |
| Privacy Settings | Per-user data sharing, community visibility, tracking consent |
| Cultural Profile | Ethnicity, religion, values, language preferences, communication style |
| Onboarding Flow | Multi-step MongoDB-backed onboarding (questions, goals, activities), separate flows for users and hosts |
| Feature | Description |
|---|---|
| Multi-modal Logging | Text label, photo emotion (DeepFace), and voice sentiment (HuBERT), combined into finalMood |
| Synchronous Analysis | Photo/voice analysis runs inline in POST /v1/mood-log, the response carries the real finalMood and matching experiences immediately, no polling or socket wait |
| Daily Summaries | Mood breakdown grouped into morning, afternoon, night |
| Streak & Heatmap | Consecutive-day streak calculation and date-to-mood heatmap data |
| File Storage | Uploaded media saved locally or to S3, downloaded back to a temp file when re-sent to FastAPI |
| Feature | Description |
|---|---|
| Race-safe Booking | Single CTE with SELECT ... FOR UPDATE inside a DB transaction atomically checks capacity and reserves a spot, see Engineering Challenges |
| Cancellation & Rebooking | Cancelling sets status='cancelled'; rebooking the same experience restores the row instead of inserting a duplicate |
| Host Dashboard | Revenue, average rating, 90-day trend, funnel, and emotional-outcome stats |
| QR Check-in | Signed token (ATTENDANCE_JWT_SECRET) turned into a QR code for in-person check-in |
| Real-time Spots | Socket.IO room per experience broadcasts spotsLeft after every booking or cancellation |
| AI Experience Generation | Host speaks or types a description; Gemini turns it into structured experience fields |
| Feature | Description |
|---|---|
| Mood-driven Matching | Active path: generateForUserByMood matches experiences to the user's latest detected mood |
| Embedding Search | Secondary path: Mongo vector search + LLM rerank (OpenAI or Gemini, RANKING_PROVIDER), implemented but not yet wired to a controller |
| Real-time Push | Recommendations delivered over Socket.IO the moment the worker finishes processing |
| Feature | Description |
|---|---|
| Group Management | Host-created communities, public search/filter, privacy settings |
| Membership Roles | member / moderator / admin per community |
| Posts, Reactions, Comments | Idempotent reaction upsert, cursor-paginated comments, author-only deletes |
| Feature | Description |
|---|---|
| Post-Experience Feedback | Rating and comment, duplicate prevention |
| Automated Reminders | Cron job finds ended experiences and enqueues a Bull job per attendee (see the note on the cron expression under Known Gaps) |
| In-app + Email + Push (partial) | Notification row and Socket.IO push always happen; email is queued through BullMQ; push is stubbed |
| Feature | Description |
|---|---|
| API Documentation | Swagger UI at /api-docs |
| Job Monitoring | Bull Board dashboard at /admin/queues |
| Live Module Graph | GET /v1/diagram renders a Mermaid dependency graph of the running app via nestjs-spelunker |
| Structured Logging | Winston, daily-rotating file transport, per-request id via AsyncLocalStorage |
| Input Validation | Global ValidationPipe + class-validator on every DTO |
1. Preventing double-booking under concurrency.
Two users tapping "book" on the last spot of an experience at the same instant is a classic race condition. BookingCreationService.tryCreateOrRestoreBooking (src/booking/services/user/booking-creation.service.ts) handles it with a single raw-SQL statement, executed inside a READ COMMITTED transaction opened by TransactionService (src/common/services/transaction.service.ts):
WITH selected_experience AS (
SELECT e.* FROM "experience" e WHERE e.id = $1 FOR UPDATE -- row-locks the experience
), ...
insert_new AS (
INSERT INTO "booking" (...)
SELECT $1, $2, 'confirmed', now()
FROM selected_experience se
WHERE NOT EXISTS (...) AND se."spotsFilled" < se."totalSpots" -- capacity check and insert, atomically
RETURNING *
), update_spots AS ( UPDATE "experience" SET "spotsFilled" = "spotsFilled" + 1 ... )The FOR UPDATE lock serializes concurrent bookings for the same experience at the database level, so the capacity check and the insert can never race: a second transaction blocks on the row lock until the first commits, then sees the updated spotsFilled and correctly fails with "No available spots". This is stronger than an application-level Redis lock because it can't drift out of sync with the actual row.
Note:
RedisServicealso exposesacquireLock/releaseLock(Lua-script-basedSET NXplus compare-and-delete), but it is not currently called anywhere in the codebase, booking concurrency is handled entirely at the Postgres layer today.
2. Choosing immediacy over decoupling for mood detection.
Emotion analysis (HuBERT, DeepFace) calls an external FastAPI service and takes real time per request. POST /v1/mood-log used to persist the raw log, return 201 immediately, and emit a mood.detect event for the worker process to analyze asynchronously, pushing the result back over Socket.IO once ready. That was replaced with a synchronous flow: the endpoint now calls the FastAPI service directly and waits, so the client gets the real finalMood and matching experiences in the same response, at the cost of the request taking as long as the slowest of the photo/voice analysis calls. Embedding generation and Gemini-based experience-field generation still run off the request path in the worker process for the same reason the async mood path originally did, see Event-Driven Architecture.
3. Two data stores for two shapes of data. Relational, highly-related domain data (users, bookings, experiences, community) lives in PostgreSQL via TypeORM, where foreign keys and transactions matter. Loosely-structured, evolving documents (onboarding answers, embedding vectors) live in MongoDB via Mongoose, where schema flexibility matters more than joins.
4. Keeping the HTTP response fast when a write has multiple side effects. Booking creation triggers attendance-record creation and a notification and a Socket.IO broadcast. All three run after the DB transaction commits, as detached (non-awaited, .catch()-guarded) calls, so a slow email or notification path never adds latency to the booking response, at the cost of those side effects being best-effort rather than transactionally guaranteed.
5. Duplicate work across a host/user/public split. experience and booking (and onboarding) expose separate controllers per audience instead of one controller with role branching inside each handler, keeping @UseGuards/@Roles declarative and unambiguous per route, at the cost of some duplication between experience.controller.ts (legacy combined controller, still registered) and the newer controllers/experience.*.controller.ts split. Worth consolidating, see Known Gaps.
The API process and the worker process are two separate Nest applications connected only by RabbitMQ, no direct function calls cross the process boundary.
Domain registry (src/infra/config/rmq.constants.ts) defines one exchange plus one durable queue per domain:
| Domain | Exchange | Queue | Routing keys |
|---|---|---|---|
| Mood | mood-exchange |
mood-tasks |
mood.detect, mood.analyzed |
| Community | community-exchange |
community-tasks |
community.embedding.generate |
| Recommendation | recommendation-exchange |
recommendation-tasks |
recommendation.generate |
| Onboarding | onboarding-exchange |
onboarding-tasks |
onboarding.completed |
| Experience | experience-exchange |
experience-tasks |
experience.generate_ai |
Producers (ClientProxy.emit(...), fire-and-forget, no reply expected): API-side services inject the relevant client (e.g. @Inject(RMQ_DOMAINS.MOOD.CLIENT)) via RmqModule.register(...) (src/infra/rmq/rmq.module.ts).
Consumers: src/worker/main.ts boots a single Nest application (WorkerModule) that opens five separate RabbitMQ connections, one per domain, each with prefetchCount: 1 (process one message at a time per domain before acking the next):
embedding.worker.ts, listens formood.analyzedandcommunity.embedding.generate: generates and stores vector embeddings in Mongo.recommendation.worker.ts, listens forrecommendation.generate: runs the recommendation engine and pushes results over Socket.IO.onboarding.worker.ts, listens foronboarding.completed: flipsuser.onboardingCompleted.experience.worker.ts, listens forexperience.generate_ai: runs Gemini experience-field generation and saves it onto theExperiencerow.
POST /v1/mood-log no longer emits mood.detect: mood detection and recommendation matching now run synchronously inside the API request (see Sample Request Flow below). Nothing publishes to the mood exchange anymore (mood.detect lost its only producer, mood.analyzed never had one), and recommendation.generate lost its only producer along with mood.detect, so embedding.worker.ts's mood-analyzed handler and all of recommendation.worker.ts are dormant. Both are left in place rather than torn out, see worker README.
Run the worker as its own process: npm run start:worker (separate from npm run start, which only runs the API).
Two message-passing systems coexist for two different jobs: RabbitMQ moves events between the API and worker processes; BullMQ/Bull moves delayed and retryable jobs within the API process itself, backed by Redis.
Two different queue libraries are in use side by side:
main.tsinstantiates queues directly with the modernbullmqpackage (for Bull Board), while the actual job processors are registered through@nestjs/bull(a wrapper around the olderbullpackage). Functionally compatible since both point at the same Redis queues, but worth knowing if you're adding a new queue, follow the@nestjs/bullpattern used by the existing processors below.
| Queue | Registered in | Producer | Processor | Purpose |
|---|---|---|---|---|
notification-queue |
src/notification/notification.module.ts |
NotificationService.createAndSend() |
src/notification/jobs/notification.processor.ts |
Sends email via Nodemailer (type: 'email'); type: 'push' is stubbed, not implemented |
feedback-request |
src/feedback/queues/feedback-queue.module.ts |
src/feedback/jobs/feedback.cron.ts (@Cron, @nestjs/schedule) |
src/feedback/queues/feedback-request.processor.ts |
Creates a PendingFeedback row per attendee once an experience's session has ended |
mood-queue |
src/infra/bull-board/bull-board.module.ts |
registered for Bull Board visibility | none | Currently monitoring-only; mood processing itself is synchronous inside POST /v1/mood-log, not queued anywhere |
All three queues are also visible (jobs, retries, failures) in Bull Board at GET /admin/queues, see src/infra/bull-board/README.md.
POST /v1/user/bookings end-to-end, tracing real files:
sequenceDiagram
participant C as Client
participant G as Guards
participant Ctrl as UserBookingController
participant Svc as BookingService
participant Create as BookingCreationService
participant DB as PostgreSQL
participant Side as BookingSideEffectsService
participant WS as ExperienceGateway
C->>G: POST /v1/user/bookings with experienceId
G->>G: Verify JWT and role
G->>Ctrl: Forward validated request
Ctrl->>Svc: createBooking(userId, dto)
Svc->>Create: createBooking(userId, dto)
Create->>DB: Begin transaction, READ COMMITTED
Create->>DB: Lock experience row, check capacity, insert or restore booking
DB-->>Create: Booking row, or a failure reason
Create->>DB: Commit transaction
Create->>Side: Queue side effects, not awaited
Create->>WS: Emit updated spot count
Create-->>Svc: Success result
Svc-->>Ctrl: Success result
Ctrl-->>C: 201 Created
Note over Side,C: Afterward, asynchronously: attendance row created,<br/>notification saved, email queued, Socket.IO push sent
Key points this illustrates:
- The guard chain runs before the DTO is even parsed by
ValidationPipe, so unauthenticated requests never reach validation or the service layer. - The entire capacity-check-and-reserve step is one database round trip inside one transaction (see Engineering Challenges), there is no read-then-write gap for a race to exploit.
- Everything after the transaction commits (attendance, notification, email, Socket.IO) is fire-and-forget: the client gets its
201as soon as the booking row exists, not after every side effect completes.
POST /v1/mood-log with a voice/photo upload, entirely within the API process, no worker hop:
sequenceDiagram
participant C as Client
participant Ctrl as MoodLogController
participant Svc as MoodLogService
participant FA as FastAPI Service
participant DB as PostgreSQL
participant Rec as ExperienceRecommendationService
C->>Ctrl: POST /v1/mood-log, multipart photo/voice plus moodLabel
Ctrl->>Svc: createForUser(userId, dto, files)
Svc->>FA: Analyze image emotion, if photo present
Svc->>FA: Analyze voice emotion, if voice present
FA-->>Svc: Dominant emotion per modality
Svc->>Svc: finalMood equals photoEmotion or voiceSentiment or moodLabel or neutral
Svc->>DB: Insert MoodLog with real finalMood
Svc->>Rec: recommendByEmotion(finalMood, userId, limit)
Rec->>DB: Match experiences by target emotion
DB-->>Rec: Matching experiences
Rec-->>Svc: Recommendations
Svc-->>Ctrl: moodLog plus recommendations
Ctrl-->>C: 201 with moodLog and recommendations
The client waits on the FastAPI call (photo/voice analysis), but gets the real finalMood and matching experiences in the same response, no polling, no Socket.IO connection needed. This used to be async, queued onto RabbitMQ for a separate worker process to handle and push back over Socket.IO, see Engineering Challenges point 2 for the tradeoff and why it changed.
Two databases, both configured in src/database/database.module.ts, shared by the API process and the worker process (each process opens its own connection pool to each):
| Store | Driver | Used for |
|---|---|---|
| PostgreSQL | TypeORM (pg) |
Users, experiences, bookings, attendance, feedback, notifications, community, anything relational |
| MongoDB | Mongoose | Onboarding documents, vector embeddings (experience/moodlog/community/post) |
synchronize: NODE_ENV !== 'production': schema auto-syncs from entities outside production; production relies on TypeORM migrations undersrc/database/migrations/(run vianpm run migration:run:prod).- Production Postgres connections use SSL with the bundled RDS CA cert (
certs/rds-ca.pem). - No explicit pool size is configured anywhere (
extra.max,poolSize, etc. are all absent), both drivers run on their library defaults. See Scale for what that means in practice.
This section describes the capacity implied by what's actually configured in this repo today, not a load-tested SLA. There is no extra.max, poolSize, PgBouncer, Node cluster mode, PM2 config, or Docker/Kubernetes replica setup anywhere in the repository, so the numbers below describe a single API instance plus a single worker instance, each talking directly to Postgres, Mongo, and Redis.
| Resource | Default in effect | Source |
|---|---|---|
| Postgres pool (API process) | max: 10 connections (node-postgres / pg default, TypeORM doesn't override it) |
src/database/database.module.ts |
| Postgres pool (worker process) | max: 10 connections, a separate pool from the API's |
src/worker/worker.module.ts (imports the same DatabaseModule, but as its own process it gets its own pool) |
| Mongoose pool (per process) | maxPoolSize: 100 (Mongoose default) |
src/database/database.module.ts |
| Redis | 1 TCP connection via ioredis, no connection pooling (Redis is single-threaded per connection anyway; BullMQ opens its own separate connections per queue) |
src/infra/redis/redis.service.ts, src/main.ts |
| RabbitMQ worker concurrency | prefetchCount: 1 times 5 domain connections = at most 5 events processed in parallel, one per domain, across the whole worker process |
src/worker/main.ts |
| HTTP concurrency (API) | Bounded by Node's single-threaded event loop per process, no cluster/PM2 config present, so one process is one CPU core doing request handling | src/main.ts |
What this means concretely:
- Concurrent HTTP connections: Node/Express can accept and hold open many hundreds of concurrent sockets fine (I/O is non-blocking), but any handler that touches Postgres is capped at 10 simultaneous in-flight queries per API instance. The 11th concurrent DB-bound request doesn't fail, it queues inside node-postgres's internal pool wait queue until a connection frees up, and only fails if the caller-side timeout (client/proxy) is hit first, since no explicit
connectionTimeoutMillisis set on the pool. - How many users, and how many concurrent users: there's no hard cap enforced by the app (rate limiting is installed but disabled, see below), so this is really "how many concurrent DB-bound requests can be serviced without queuing," which today is around 10 per API instance. Read-heavy, non-DB-bound traffic (like static Swagger docs) isn't affected by this ceiling. Realistic total concurrent users the system can serve without visibly increasing latency depends on how many requests per user are in flight at once, but with a 10-connection pool and typical query times in the tens of milliseconds, sustained throughput is on the order of a few hundred DB-touching requests per second on default hardware before the pool becomes the bottleneck, well below what the underlying Postgres instance itself could otherwise support.
- Where overflow goes: it doesn't get rejected, it queues, at two levels: (1) inside the node-postgres pool's wait queue once more than 10 queries are in flight, and (2) inside RabbitMQ's durable queues once event throughput exceeds what
prefetchCount: 1per domain can drain (messages simply sit in the queue, durable, so they survive a worker restart, until a worker connection is free to consume the next one). Neither layer drops work; both simply add latency under load. There is currently no dead-letter queue, no queue-depth alerting, and no backpressure signal sent back to the API when the worker falls behind. - Max simultaneous DB connections across the whole system: with one API instance plus one worker instance, that's
10 (API) + 10 (worker) = 20Postgres connections at steady state, plus a handful more from ad-hoctypeormCLI usage (migrations). A typical small managed Postgres tier (e.g. AWS RDSdb.t3.micro) capsmax_connectionsaround 65 to 100, so this repo's defaults leave headroom for roughly 3 to 4 more full API+worker replica pairs before hitting that ceiling, but nothing here coordinates that; adding replicas today would require either raisingmax_connections, adding an explicitextra.maxper instance, or introducing a pooler (PgBouncer / RDS Proxy) so pool sizes don't need to shrink as replicas grow. - Rate limiting is not currently active.
@nestjs/throttleris installed and pre-configured (default: 120 req/min,auth: 15 req/min,login: 5 req/min) but the wholeThrottlerModule.forRoot([...])block and itsAPP_GUARDprovider are commented out insrc/app.module.ts. Several controllers still carry@SkipThrottle(), which is currently a no-op since noThrottlerGuardis registered anywhere. Until this is re-enabled, nothing in the app itself protects against a single client opening far more concurrent requests than the pool can serve.
To raise these ceilings, in order of effort: (1) set extra: { max: N } on the TypeORM config and re-tune per available max_connections; (2) re-enable ThrottlerModule in app.module.ts; (3) run the API behind a load balancer with multiple replicas plus a connection pooler in front of Postgres; (4) raise prefetchCount on the worker's RMQ connections (trades ordering/backpressure safety for throughput); (5) move Redis to a clustered/managed tier if BullMQ throughput becomes the bottleneck.
- JWT access and refresh tokens (
@nestjs/jwt), both currently signed with a 7-day expiry (src/auth/auth.service.ts), the refresh token hash is stored (bcrypt) on theUserrow and rotated on every/auth/refreshcall. - The
jwthttpOnly cookie itself is set withmaxAge: 24h(src/auth/auth.controller.ts), which is shorter than the JWT's own 7-day expiry, the cookie expires and forces a re-login well before the token would. - Google OAuth 2.0 via Passport (
src/auth/strategies/google.strategy.ts): find-or-create by email, then issues the same JWT pair. - Guards:
JwtCookieGuard(cookie, falls back toAuthorization: Bearer),JwtBearerGuard(bearer-only),RolesGuard+@Roles()(src/common/roles.guard.ts). Most protected routes stack all three, thoughJwtBearerGuardis redundant whereverJwtCookieGuardis also present. - Attendance uses a separate signed token (
ATTENDANCE_JWT_SECRET/ATTENDANCE_JWT_EXPIRATION) for QR check-in, independent of the login JWT. - Full endpoint reference: src/auth/README.md.
This NestJS backend calls out to a separate Python FastAPI service (not part of this repository) via ApiClientService (src/common/services/api-client.service.ts, base URL FASTAPI_URL, bearer-token'd with HF_TOKEN). It exposes three endpoints:
| Endpoint | Model | Purpose |
|---|---|---|
POST /analyze-voice-emotion |
HuBERT (superb/hubert-large-superb-er) |
WAV upload to dominant emotion plus per-class confidence |
POST /analyze-image-emotion |
DeepFace | Image upload to dominant facial emotion |
POST /embed |
SentenceTransformer (all-MiniLM-L6-v2) |
Text to a 384-dimensional embedding vector |
All three models load once at FastAPI startup, not per request, to keep inference latency predictable. See src/mood-log/services/emotion-analysis.service.ts and src/embedding/services/embedding.service.ts for the calling code on this side.
See .env.example for the full list of keys this repo reads (no real values are committed, .env is gitignored). Grouped summary:
| Group | Keys |
|---|---|
| Postgres | POSTGRES_HOST, POSTGRES_PORT, POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB |
| Mongo | MONGO_URI, MONGO_DB |
| Redis | REDIS_URL |
| App | NODE_ENV, PORT, FRONTEND_URL |
| Auth | JWT_SECRET, JWT_REFRESH_SECRET, ATTENDANCE_JWT_SECRET, ATTENDANCE_JWT_EXPIRATION, GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET, GOOGLE_CALLBACK_URL |
| RabbitMQ | RABBITMQ_URL, RMQ_MOOD_QUEUE/RMQ_MOOD_EXCHANGE, RMQ_COMM_QUEUE/RMQ_COMM_EXCHANGE, RMQ_REC_QUEUE/RMQ_REC_EXCHANGE, RMQ_ONBOARDING_QUEUE/RMQ_ONBOARDING_EXCHANGE, RMQ_EXP_QUEUE/RMQ_EXP_EXCHANGE |
| AI | GEMINI_API_KEY, OPENAI_API_KEY, RANKING_PROVIDER, FASTAPI_URL, HF_TOKEN |
| Storage | AWS_REGION, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, S3_BUCKET, S3_BASE_URL, EXPERIENCE_IMAGES_CDN_URL, UPLOADS_DIR |
SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASS, SMTP_FROM |
|
| Logging | LOG_LEVEL, LOG_DIR |
| Observability | OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_HEADERS, OTEL_SERVICE_NAME, LOKI_URL, LOKI_USER, LOKI_API_KEY, see Observability |
Double-check
RMQ_REC_QUEUE/RMQ_REC_EXCHANGEagainst your deployment's.env,src/infra/config/rmq.constants.tsreads those exact names, while some.envfiles in the wild useRMQ_RECOMMENDATION_QUEUEinstead, which the code will silently ignore in favor of the hardcoded default.
# install dependencies
npm install
# build (verifies the whole project compiles)
npm run build
# run the unit test suite (no external services required, everything is mocked)
npm test
# start the API server, http://localhost:3002, Swagger at /api-docs, Bull Board at /admin/queues
npm run start
# start the worker process, consumes RabbitMQ, must run alongside the API for
# mood analysis, recommendations, and AI experience generation to complete
npm run start:worker
# generate a new TypeORM migration after changing an entity
npx typeorm migration:generate -n MigrationName
# apply migrations in production
npm run migration:run:prodRequires a reachable PostgreSQL instance, MongoDB instance, Redis instance, and RabbitMQ broker, see .env.example for every variable that needs a value. The quickest way to get the whole stack running locally, including Postgres/Mongo/Redis/RabbitMQ, is docker compose up --build, see Deployment below.
Both processes build from a single Dockerfile (multi-stage: compiles once with nest build, then a slim runtime image; CMD runs the API, the worker overrides command: to run dist/worker/main.js instead), since they share the same src/ tree and nest build compiles the whole thing regardless of entrypoint.
Local, full stack: docker-compose.yml runs Postgres, MongoDB, Redis, RabbitMQ (with its management UI on :15672), plus the api and worker services, wired with health checks so the app waits for its dependencies to actually be ready. It loads your .env for app-level secrets, then overrides the 4 infra connection vars to point at the containers on the compose network instead of whatever's in .env, and forces NODE_ENV=development so synchronize creates tables automatically against the throwaway local Postgres volume.
docker compose up --buildCloud deploy: the same image is deploy-ready for Fly.io, a plain VPS, or ECS. The 4 infra dependencies (Postgres, MongoDB, Redis, RabbitMQ) need to be reachable over the public internet from wherever compute runs, free options that work well here: Neon (Postgres), MongoDB Atlas M0, Upstash (Redis, use the TCP/rediss:// connection string, not the REST API URL), and CloudAMQP (RabbitMQ). Set NODE_ENV=production so TypeORM SSL and synchronize: false kick in, then run npm run migration:run:prod once against the target database before the app's first boot.
Logging already goes through nest-winston/winston (console + local DailyRotateFile, see src/main.ts/src/worker/main.ts). On top of that:
- Traces + metrics:
src/tracing.tssets up the OpenTelemetry Node SDK with auto-instrumentation (HTTP,pg,mongoose,ioredis,amqplib), exporting over OTLP. It's preloaded vianode -r ./dist/tracing.js dist/main.js(see thestart/start:workerscripts, theDockerfileCMD, anddocker-compose.yml), not imported normally, so instrumentation patches modules before the app itself first requires them. Entirely opt-in: withOTEL_EXPORTER_OTLP_ENDPOINTunset it logs a line and skips setup, nothing else changes. - Logs shipping: a
winston-lokitransport is added alongside the existing console/file transports whenLOKI_URLis set, labeled per-process (OTEL_SERVICE_NAME, e.g.moodly-api/moodly-worker) so both processes' logs are distinguishable in one place. - Destination: Grafana Cloud free tier (Loki for logs, Mimir for metrics, Tempo for traces, one account, one set of credentials). See
.env.examplefor the exact 6 variables needed. - Alerting: not wired into app code, Grafana Cloud's Alerting has native Slack and Discord contact points (webhook-based, no OAuth app needed for either), configured directly in its UI once logs/metrics are flowing. Not yet set up for this project.
| Item | Notes |
|---|---|
| Rate limiting disabled | ThrottlerModule is fully configured but commented out in src/app.module.ts, see Scale |
| Inconsistent bcrypt cost factors | Password hashing uses cost 12, refresh-token hashing uses cost 10 (src/auth/auth.constants.ts), not a documented design choice, worth deciding whether to unify |
No start:dev npm script |
@nestjs/cli is a dependency but there's no watch-mode script; run npx nest start --watch directly for now |
| No end-to-end tests yet | npm run test:e2e is wired to test/jest-e2e.json and passes trivially (--passWithNoTests), but no *.e2e-spec.ts files exist yet, only unit tests (npm test) currently exercise real behavior |
| Legacy duplicate experience controller | src/experience/experience.controller.ts overlaps with the newer controllers/experience.host/public/user.controller.ts split, both are currently registered |
POST /v1/notification role check is inert |
@Roles('host') is set on the handler but RolesGuard isn't in that route's @UseGuards(...), so the role restriction doesn't actually run, see src/notification/README.md |
| Feedback reminder cron expression | @Cron('0 */19999 * * * *') in src/feedback/jobs/feedback.cron.ts doesn't match its "every 5 min" comment, worth re-checking before relying on it |
| Embedding-based recommendation path unused | RecommendationService.generateForUser() (embedding + LLM rerank) exists and works but nothing currently calls it, only the mood-based path is wired to a controller/worker |
| Grafana alerting not configured | OTel/Loki wiring exists (see Observability) and Grafana Cloud supports Slack/Discord contact points natively, but no alert rules or contact points have been created for this project yet |
| Hybrid Recommendation Engine | Planned: merge emotion-based and embedding-based results with scored deduplication |
| Participant Matchmaking / Engagement Analytics | idealParticipantTraits and engagementStats fields already exist on Experience for this, not yet consumed |
| Extended RMQ domains | Architecture supports adding FEEDBACK/BOOKINGS domains the same way the existing five were added |