Skip to main content
Back to Blog
Payment Gateway IntegrationNode.jsExpress.jsPostgreSQLStripeRazorpayCashfreen8nWebhooksIdempotencySocket.ioAWSDockerCI/CDSaaS Billing

Architecting Multi-Gateway Payment Systems: Orchestrating Webhooks with Node.js, n8n, and PostgreSQL on AWS

A practical 2026 blueprint for building a resilient multi-gateway payment backend. Integrate Stripe, Razorpay, and Cashfree with Express.js and PostgreSQL idempotency, automate reconciliation with self-hosted n8n and Make.com, push real-time updates with Socket.io, and ship it all with Docker, CI/CD, and AWS.

October 8, 202624 min readNiraj Kumar

Architecting Multi-Gateway Payment Systems: Orchestrating Webhooks with Node.js, n8n, and PostgreSQL on AWS

If you have ever been paged at 2 a.m. because a customer was charged twice, or worse, charged once and never given access, you already know that payments are not a feature. They are a distributed system wearing a feature's clothing.

The moment you add a second gateway, the problem multiplies. Stripe signs webhooks one way, Razorpay another, and Cashfree a third. Amounts arrive in different units. Event names do not match. Retries come out of order. Somebody on your team writes a cron job to "fix the stuck ones," and six months later nobody remembers what it does.

This guide is the blueprint I wish I had when I first built a cross-border billing backend. We will design a payment platform that supports Stripe, Razorpay, and Cashfree behind a single internal contract, using:

  • Node.js 24 LTS and Express.js for the REST API and webhook receiver
  • PostgreSQL for transactional state, idempotency, and the webhook inbox
  • Socket.io for real-time payment status in the browser
  • Self-hosted n8n (plus Make.com for business teams) to replace expensive, brittle custom cron jobs with visible, retryable reconciliation workflows
  • Docker, GitHub Actions, and AWS (ECS Fargate, RDS, ElastiCache) for a production-grade delivery pipeline

By the end you will have working code patterns, a deployment shape you can copy, and a list of mistakes to avoid before they cost you money.

High-level architecture of a multi-gateway payment system with Express.js, PostgreSQL, n8n, Socket.io and AWS

Why Multi-Gateway Is the Default in 2026

A single gateway used to be a sensible default. In 2026 it is a risk you have to justify. The reasons are practical rather than fashionable:

  • Payment method coverage. UPI, netbanking, and wallets matter enormously in India, which is where Razorpay and Cashfree shine. Stripe remains the strongest option for international cards, subscriptions, and tax tooling.
  • Resilience. Gateways have incidents. If your checkout can fall back to a second provider for new payments, a provider outage becomes an annoyance instead of an revenue event.
  • Cost and negotiation. Routing by currency, country, or payment method lets you optimize fees and gives you leverage at renewal time.
  • Regulation and data residency. Some markets prefer or require local processors.

The catch is that every additional gateway multiplies integration surface area. The solution is not to write three integrations. It is to write one internal payment model and make each gateway conform to it.

Architecture at a Glance

Here is the shape of the system we are building:

  1. Client creates a payment through POST /v1/payments with an Idempotency-Key header.
  2. API inserts a payments row, picks a gateway using routing rules, and returns checkout details.
  3. Gateway collects the money and sends webhooks to POST /webhooks/:gateway.
  4. Webhook receiver verifies the signature, stores the raw event in webhook_events, and returns 200 quickly.
  5. Worker claims events from the inbox, normalizes them, and applies state changes inside a transaction.
  6. Socket.io (via a Redis emitter) tells the user's browser that something changed.
  7. n8n runs scheduled reconciliation, retries stuck payments, and alerts humans when drift appears.
  8. Make.com handles the finance-team side: spreadsheets, tickets, and summaries.

The key design principle is the inbox pattern: the HTTP layer only accepts events, and a separate process handles them. That one decision removes an entire category of bugs.

Start with the Data Model: Idempotency Lives in the Database

Application-level checks like "have I seen this event?" fail under concurrency. Two webhook deliveries can arrive within milliseconds of each other on different containers. The only reliable referee is a database constraint.

-- 001_init.sql
CREATE EXTENSION IF NOT EXISTS pgcrypto;

CREATE TABLE payments (
  id                 UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  user_id            UUID        NOT NULL,
  gateway            TEXT        NOT NULL CHECK (gateway IN ('stripe', 'razorpay', 'cashfree')),
  gateway_payment_id TEXT,
  idempotency_key    TEXT        NOT NULL UNIQUE,
  amount_minor       BIGINT      NOT NULL CHECK (amount_minor > 0),
  currency           CHAR(3)     NOT NULL,
  status             TEXT        NOT NULL DEFAULT 'pending'
                     CHECK (status IN ('pending', 'failed', 'succeeded', 'refunded')),
  created_at         TIMESTAMPTZ NOT NULL DEFAULT now(),
  updated_at         TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE UNIQUE INDEX payments_gateway_payment_uidx
  ON payments (gateway, gateway_payment_id)
  WHERE gateway_payment_id IS NOT NULL;

CREATE TABLE webhook_events (
  id              BIGSERIAL PRIMARY KEY,
  gateway         TEXT        NOT NULL,
  event_id        TEXT        NOT NULL,
  event_type      TEXT        NOT NULL,
  payload         JSONB       NOT NULL,
  status          TEXT        NOT NULL DEFAULT 'received'
                  CHECK (status IN ('received', 'retry', 'processed', 'dead')),
  attempts        INT         NOT NULL DEFAULT 0,
  next_attempt_at TIMESTAMPTZ NOT NULL DEFAULT now(),
  last_error      TEXT,
  received_at     TIMESTAMPTZ NOT NULL DEFAULT now(),
  processed_at    TIMESTAMPTZ,
  UNIQUE (gateway, event_id)
);

CREATE INDEX webhook_events_pending_idx
  ON webhook_events (next_attempt_at)
  WHERE status IN ('received', 'retry');

-- Forward-only state machine: a late "failed" must never overwrite "succeeded".
CREATE FUNCTION payment_status_rank(s TEXT) RETURNS INT
LANGUAGE sql IMMUTABLE AS $$
  SELECT CASE s
    WHEN 'pending'   THEN 0
    WHEN 'failed'    THEN 1
    WHEN 'succeeded' THEN 2
    WHEN 'refunded'  THEN 3
  END
$$;

Three details worth calling out:

  • Money is stored as integer minor units (amount_minor). Floats and money should never meet.
  • UNIQUE (gateway, event_id) is your webhook deduplication. Not a cache, not a Redis key. A constraint.
  • payment_status_rank gives you a tiny state machine in SQL so out-of-order events cannot move a payment backwards.

The Gateway Adapter Pattern

Each gateway gets an adapter that implements the same small interface:

  • verify(rawBody, headers) returns the verified event ID, type, and parsed payload (or throws)
  • normalize(payload) maps the vendor event to our internal event shape (or returns null if we do not care)
  • createCheckout(payment) starts a payment on that gateway

Internally, we only ever speak this normalized shape:

// { type: 'payment.succeeded' | 'payment.failed', paymentRef: <our payments.id>,
//   gatewayPaymentId: string, amountMinor: number, currency: string }

To make this work, we pass our own payments.id to every gateway (as Stripe metadata, Razorpay notes, or the Cashfree order_id). That gives us a stable join key regardless of vendor.

Stripe

Stripe's SDK does the heavy lifting, but note that it needs the raw request body, not the parsed JSON.

// src/gateways/stripe.js
import Stripe from 'stripe';

const stripe = new Stripe(process.env.STRIPE_SECRET_KEY);

const TYPE_MAP = {
  'payment_intent.succeeded': 'payment.succeeded',
  'payment_intent.payment_failed': 'payment.failed',
};

export const stripeAdapter = {
  name: 'stripe',

  verify(rawBody, headers) {
    const event = stripe.webhooks.constructEvent(
      rawBody,
      headers['stripe-signature'],
      process.env.STRIPE_WEBHOOK_SECRET
    );
    return { eventId: event.id, type: event.type, payload: event };
  },

  normalize(event) {
    const type = TYPE_MAP[event.type];
    if (!type) return null;
    const pi = event.data.object;
    return {
      type,
      paymentRef: pi.metadata.payment_id,
      gatewayPaymentId: pi.id,
      amountMinor: pi.amount,
      currency: pi.currency.toUpperCase(),
    };
  },

  async createCheckout(payment) {
    const intent = await stripe.paymentIntents.create(
      {
        amount: Number(payment.amount_minor),
        currency: payment.currency.toLowerCase(),
        automatic_payment_methods: { enabled: true },
        metadata: { payment_id: payment.id },
      },
      { idempotencyKey: payment.idempotency_key }
    );
    return { gatewayPaymentId: intent.id, clientSecret: intent.client_secret };
  },
};

Stripe honors idempotency keys natively on write requests, so we pass ours straight through.

Razorpay

Razorpay signs the raw body with HMAC-SHA256 using your webhook secret and sends it in X-Razorpay-Signature. It also sends a unique X-Razorpay-Event-Id, which is exactly what we need for deduplication.

// src/gateways/razorpay.js
import crypto from 'node:crypto';

const safeEqual = (a, b) => {
  const ab = Buffer.from(a);
  const bb = Buffer.from(b);
  return ab.length === bb.length && crypto.timingSafeEqual(ab, bb);
};

export const razorpayAdapter = {
  name: 'razorpay',

  verify(rawBody, headers) {
    const expected = crypto
      .createHmac('sha256', process.env.RAZORPAY_WEBHOOK_SECRET)
      .update(rawBody)
      .digest('hex');

    if (!safeEqual(expected, headers['x-razorpay-signature'] ?? '')) {
      throw new Error('Invalid Razorpay signature');
    }

    const payload = JSON.parse(rawBody.toString('utf8'));
    const eventId =
      headers['x-razorpay-event-id'] ??
      crypto.createHash('sha256').update(rawBody).digest('hex');

    return { eventId, type: payload.event, payload };
  },

  normalize(payload) {
    const map = { 'payment.captured': 'payment.succeeded', 'payment.failed': 'payment.failed' };
    const type = map[payload.event];
    if (!type) return null;
    const p = payload.payload.payment.entity;
    return {
      type,
      paymentRef: p.notes?.payment_id,
      gatewayPaymentId: p.id,
      amountMinor: p.amount, // already in paise
      currency: p.currency,
    };
  },
};

Cashfree

Cashfree's webhook signature is computed over the timestamp header concatenated with the raw body, using your secret key, then Base64-encoded. Its amounts arrive as decimal numbers in major units, which is a classic trap.

// src/gateways/cashfree.js
import crypto from 'node:crypto';

const toMinor = (value) => {
  // Parse from string to avoid float drift: "499.50" -> 49950
  const [whole, frac = ''] = String(value).split('.');
  return Number(whole) * 100 + Number(frac.padEnd(2, '0').slice(0, 2));
};

export const cashfreeAdapter = {
  name: 'cashfree',

  verify(rawBody, headers) {
    const timestamp = headers['x-webhook-timestamp'] ?? '';
    const received = headers['x-webhook-signature'] ?? '';
    const expected = crypto
      .createHmac('sha256', process.env.CASHFREE_SECRET_KEY)
      .update(timestamp + rawBody.toString('utf8'))
      .digest('base64');

    if (expected !== received) throw new Error('Invalid Cashfree signature');

    const payload = JSON.parse(rawBody.toString('utf8'));
    // No dedicated event ID header: derive a deterministic one.
    const eventId = crypto.createHash('sha256').update(timestamp + rawBody).digest('hex');
    return { eventId, type: payload.type, payload };
  },

  normalize(payload) {
    const map = {
      PAYMENT_SUCCESS_WEBHOOK: 'payment.succeeded',
      PAYMENT_FAILED_WEBHOOK: 'payment.failed',
    };
    const type = map[payload.type];
    if (!type) return null;
    const { order, payment } = payload.data;
    return {
      type,
      paymentRef: order.order_id, // we set order_id = payments.id
      gatewayPaymentId: String(payment.cf_payment_id),
      amountMinor: toMinor(payment.payment_amount),
      currency: payment.payment_currency,
    };
  },
};

Always check the vendor's current webhook documentation before shipping, since header names and payload shapes evolve. The pattern, however, stays the same: verify, derive a stable event ID, normalize.

The Webhook Receiver: Verify, Persist, Acknowledge

The receiver has exactly three jobs. Notice how little it does.

// src/routes/webhooks.js
import express from 'express';
import { adapters } from '../gateways/index.js';
import { pool } from '../db.js';

export const webhookRouter = express.Router();

// express.raw MUST be used here. Mount this router BEFORE any global express.json().
webhookRouter.post(
  '/:gateway',
  express.raw({ type: '*/*', limit: '1mb' }),
  async (req, res) => {
    const adapter = adapters[req.params.gateway];
    if (!adapter) return res.sendStatus(404);

    let verified;
    try {
      verified = adapter.verify(req.body, req.headers);
    } catch (err) {
      req.log.warn({ gateway: adapter.name, err: err.message }, 'webhook rejected');
      return res.sendStatus(400);
    }

    try {
      await pool.query(
        `INSERT INTO webhook_events (gateway, event_id, event_type, payload)
         VALUES ($1, $2, $3, $4)
         ON CONFLICT (gateway, event_id) DO NOTHING`,
        [adapter.name, verified.eventId, verified.type, verified.payload]
      );
      return res.sendStatus(200); // duplicates are also a success from the gateway's view
    } catch (err) {
      req.log.error({ err }, 'failed to persist webhook');
      return res.sendStatus(500); // let the gateway retry
    }
  }
);

A duplicate delivery hits ON CONFLICT DO NOTHING, does nothing, and still returns 200. The gateway is happy, and your data stays clean.

Sequence diagram showing webhook verification, inbox insert, worker processing and Socket.io notification

The Worker: Process the Inbox Safely

PostgreSQL gives us a surprisingly good job queue with FOR UPDATE SKIP LOCKED. Multiple workers can pull events concurrently without ever grabbing the same row, and you do not need to introduce another queueing system on day one.

// src/worker/process.js
import { pool } from '../db.js';
import { adapters } from '../gateways/index.js';
import { emitter } from '../realtime/emitter.js';

export async function processNext() {
  const client = await pool.connect();
  let evt;
  try {
    await client.query('BEGIN');

    const { rows } = await client.query(
      `SELECT * FROM webhook_events
       WHERE status IN ('received', 'retry') AND next_attempt_at <= now()
       ORDER BY id
       LIMIT 1
       FOR UPDATE SKIP LOCKED`
    );
    evt = rows[0];
    if (!evt) {
      await client.query('COMMIT');
      return false;
    }

    const normalized = adapters[evt.gateway].normalize(evt.payload);
    let updated = null;

    if (normalized) {
      const res = await client.query(
        `UPDATE payments
            SET status = $2,
                gateway_payment_id = COALESCE(gateway_payment_id, $3),
                updated_at = now()
          WHERE id = $1
            AND payment_status_rank(status) < payment_status_rank($2)
        RETURNING id, user_id, status`,
        [
          normalized.paymentRef,
          normalized.type === 'payment.succeeded' ? 'succeeded' : 'failed',
          normalized.gatewayPaymentId,
        ]
      );
      updated = res.rows[0] ?? null;
    }

    await client.query(
      `UPDATE webhook_events SET status = 'processed', processed_at = now() WHERE id = $1`,
      [evt.id]
    );
    await client.query('COMMIT');

    // Notify only after the transaction is durable.
    if (updated) {
      emitter
        .to(`user:${updated.user_id}`)
        .emit('payment:updated', { id: updated.id, status: updated.status });
    }
    return true;
  } catch (err) {
    await client.query('ROLLBACK');
    if (evt) await markFailed(evt.id, err);
    return true; // keep draining the queue
  } finally {
    client.release();
  }
}

async function markFailed(id, err) {
  // Exponential backoff: 30s, 60s, 120s ... and park it as "dead" after 8 attempts.
  await pool.query(
    `UPDATE webhook_events
        SET attempts = attempts + 1,
            last_error = $2,
            status = CASE WHEN attempts + 1 >= 8 THEN 'dead' ELSE 'retry' END,
            next_attempt_at = now() + (interval '30 seconds' * power(2, attempts))
      WHERE id = $1`,
    [id, String(err.message).slice(0, 500)]
  );
}

Look closely at the UPDATE. The payment_status_rank guard means that if a delayed payment.failed arrives after payment.succeeded, the row simply does not match and nothing changes. You get idempotent, order-independent processing without writing a single if statement in JavaScript.

Creating Payments Idempotently, with Routing and Failover

On the way in, the same principle applies. The client sends an Idempotency-Key; we let PostgreSQL decide whether this is a new request or a retry.

// src/services/createPayment.js
import { pool } from '../db.js';
import { adapters } from '../gateways/index.js';

const ROUTES = { INR: ['razorpay', 'cashfree'], DEFAULT: ['stripe'] };

export async function createPayment({ userId, amountMinor, currency, idempotencyKey }) {
  const { rows } = await pool.query(
    `INSERT INTO payments (user_id, gateway, idempotency_key, amount_minor, currency)
     VALUES ($1, $2, $3, $4, $5)
     ON CONFLICT (idempotency_key) DO NOTHING
     RETURNING *`,
    [userId, (ROUTES[currency] ?? ROUTES.DEFAULT)[0], idempotencyKey, amountMinor, currency]
  );

  // Retry of an earlier request: return the existing payment instead of creating another.
  if (rows.length === 0) {
    const existing = await pool.query(
      'SELECT * FROM payments WHERE idempotency_key = $1',
      [idempotencyKey]
    );
    return { payment: existing.rows[0], replayed: true };
  }

  const payment = rows[0];
  const candidates = ROUTES[currency] ?? ROUTES.DEFAULT;
  let lastError;

  for (const gateway of candidates) {
    try {
      const checkout = await adapters[gateway].createCheckout(payment);
      await pool.query(
        `UPDATE payments SET gateway = $2, gateway_payment_id = $3 WHERE id = $1`,
        [payment.id, gateway, checkout.gatewayPaymentId]
      );
      return { payment: { ...payment, gateway }, checkout, replayed: false };
    } catch (err) {
      lastError = err;
      // Only fail over on clear, pre-creation errors (5xx, connection refused).
      // A timeout is ambiguous: the gateway may have created the order anyway.
      if (!isSafeToFailOver(err)) break;
    }
  }
  throw lastError;
}

The comment on timeouts is the part people miss. If a gateway call times out, you genuinely do not know whether the order exists. Failing over blindly can create two live checkouts for one purchase. Treat ambiguous failures as "unknown," and let reconciliation sort them out. Wrapping gateway calls in a circuit breaker (for example, opossum) also stops a degraded provider from dragging your API latency down.

Real-Time Updates with Socket.io

Users staring at a spinner after paying is a support ticket waiting to happen. Socket.io lets us push the result the moment the worker commits it.

Because the API and worker run as separate containers, the worker does not host a socket server. It uses the Redis emitter to publish into the same adapter your API instances share.

// src/realtime/server.js (runs inside the API)
import { Server } from 'socket.io';
import { createAdapter } from '@socket.io/redis-adapter';
import { createClient } from 'redis';
import jwt from 'jsonwebtoken';

export async function attachSocket(httpServer) {
  const pub = createClient({ url: process.env.REDIS_URL });
  const sub = pub.duplicate();
  await Promise.all([pub.connect(), sub.connect()]);

  const io = new Server(httpServer, {
    cors: { origin: process.env.APP_ORIGIN, credentials: true },
    adapter: createAdapter(pub, sub),
  });

  io.use((socket, next) => {
    try {
      const claims = jwt.verify(socket.handshake.auth.token, process.env.JWT_PUBLIC_KEY);
      socket.data.userId = claims.sub;
      next();
    } catch {
      next(new Error('unauthorized'));
    }
  });

  io.on('connection', (socket) => socket.join(`user:${socket.data.userId}`));
  return io;
}
// src/realtime/emitter.js (used by the worker)
import { Emitter } from '@socket.io/redis-emitter';
import { createClient } from 'redis';

const client = createClient({ url: process.env.REDIS_URL });
await client.connect();
export const emitter = new Emitter(client);

On the front end, listen for payment:updated, but also refetch the payment from the REST API when the socket connects or reconnects. Sockets are best-effort. The database is the truth.

Replacing Cron Jobs with n8n

Here is the uncomfortable truth about custom cron scripts: they work, they get forgotten, and they fail silently. When one breaks, nobody knows until finance notices missing revenue.

Self-hosted n8n gives you the same scheduling power with a canvas that shows exactly what ran, what failed, and why. Execution history, retries, error workflows, and branching come built in.

Design rule: n8n orchestrates, your API decides

I strongly recommend not putting gateway secrets or business logic inside n8n. Instead, expose two small internal endpoints (private network only, protected by an API key or mTLS):

  • GET /internal/reconcile/candidates returns payments that are pending for more than 15 minutes, or recently succeeded payments to double-check.
  • POST /internal/payments/:id/resync fetches the live status from the gateway and inserts a synthetic event into webhook_events.

That second endpoint is the clever bit. A resync does not update the payment directly. It feeds the same inbox, so reconciliation and live webhooks share one code path, one state machine, and one set of tests.

The reconciliation workflow

A simple version has these nodes:

  1. Schedule Trigger runs every 15 minutes (plus a deeper nightly run).
  2. HTTP Request calls the candidates endpoint.
  3. Split In Batches processes 50 payments at a time to be kind to gateway rate limits.
  4. HTTP Request calls the resync endpoint, with retry on failure enabled.
  5. Code aggregates the results into a drift summary.
  6. IF checks whether any drift or repeated failures were found.
  7. Slack / Email alerts the on-call engineer, and a Webhook node pings a Make.com scenario.

The aggregation step is tiny:

// n8n Code node: summarize resync results
const results = $input.all().map((i) => i.json);

const drifted = results.filter((r) => r.statusChanged);
const failed = results.filter((r) => r.error);

return [
  {
    json: {
      checked: results.length,
      driftCount: drifted.length,
      failedCount: failed.length,
      sample: drifted.slice(0, 10).map((r) => r.paymentId),
    },
  },
];

Add a second workflow with an Error Trigger and a daily Postgres query for webhook_events where status = 'dead'. Dead letters should be loud, never silent.

Where Make.com earns its place

n8n is the engineering team's tool. Finance and support teams often prefer something they can own without touching infrastructure, and Make.com fits that well. A typical scenario: receive the drift summary from n8n through a webhook, append rows to a shared spreadsheet, create a ticket for each mismatch, and post a weekly digest. The boundary stays clean: n8n runs the system, Make.com serves the business. If you do not need the business-side layer, skip Make.com entirely and keep everything in n8n.

One honest note on licensing: n8n is source-available under its own fair-code license, and self-hosting for your own internal workflows is generally fine. Read the current terms before embedding it in a product you resell.

Containerizing with Docker

A multi-stage build keeps images small and avoids shipping dev dependencies to production. The same image runs as the API or the worker, just with a different command.

# Dockerfile
FROM node:24-alpine AS deps
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev

FROM node:24-alpine AS runtime
ENV NODE_ENV=production
WORKDIR /app
RUN addgroup -S app && adduser -S app -G app
COPY --from=deps /app/node_modules ./node_modules
COPY src ./src
COPY package*.json ./
USER app
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=3s \
  CMD wget -qO- http://localhost:3000/healthz || exit 1
CMD ["node", "src/server.js"]

For local development, Docker Compose brings up the whole stack, including n8n backed by PostgreSQL instead of the default SQLite:

# docker-compose.yml
services:
  postgres:
    image: postgres:17
    environment:
      POSTGRES_PASSWORD: devpassword
      POSTGRES_DB: payments
    volumes: [pgdata:/var/lib/postgresql/data]

  redis:
    image: redis:7-alpine

  api:
    build: .
    env_file: .env
    ports: ["3000:3000"]
    depends_on: [postgres, redis]

  worker:
    build: .
    command: ["node", "src/worker/index.js"]
    env_file: .env
    depends_on: [postgres, redis]

  n8n:
    image: n8nio/n8n:latest
    ports: ["5678:5678"]
    environment:
      DB_TYPE: postgresdb
      DB_POSTGRESDB_HOST: postgres
      DB_POSTGRESDB_DATABASE: payments
      DB_POSTGRESDB_USER: postgres
      DB_POSTGRESDB_PASSWORD: devpassword
      N8N_ENCRYPTION_KEY: change-me-in-real-life
      EXECUTIONS_DATA_PRUNE: "true"
    depends_on: [postgres]

volumes:
  pgdata:

In production, pin n8n to a specific version tag rather than latest, and give it its own database schema or a dedicated database.

CI/CD: From Pull Request to Production

A payment system deserves a pipeline that refuses to ship broken code. The GitHub Actions workflow below runs tests against a real PostgreSQL service (not mocks, since your idempotency guarantees live in the database), builds the image, and deploys to ECS using short-lived OIDC credentials. No long-lived AWS keys are stored anywhere.

# .github/workflows/deploy.yml
name: ci-cd
on:
  push:
    branches: [main]
  pull_request:

permissions:
  id-token: write
  contents: read

jobs:
  test:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:17
        env:
          POSTGRES_PASSWORD: test
        ports: ["5432:5432"]
        options: >-
          --health-cmd "pg_isready -U postgres"
          --health-interval 5s --health-retries 10
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 24, cache: npm }
      - run: npm ci
      - run: npm run migrate
        env: { DATABASE_URL: postgres://postgres:test@localhost:5432/postgres }
      - run: npm test
        env: { DATABASE_URL: postgres://postgres:test@localhost:5432/postgres }

  deploy:
    needs: test
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::123456789012:role/gha-deploy
          aws-region: ap-south-1
      - id: ecr
        uses: aws-actions/amazon-ecr-login@v2
      - name: Build and push
        run: |
          IMAGE=${{ steps.ecr.outputs.registry }}/payments:${{ github.sha }}
          docker build -t $IMAGE .
          docker push $IMAGE
      - id: render
        uses: aws-actions/amazon-ecs-render-task-definition@v1
        with:
          task-definition: infra/api-task-def.json
          container-name: api
          image: ${{ steps.ecr.outputs.registry }}/payments:${{ github.sha }}
      - uses: aws-actions/amazon-ecs-deploy-task-definition@v2
        with:
          task-definition: ${{ steps.render.outputs.task-definition }}
          service: payments-api
          cluster: payments-prod
          wait-for-service-stability: true

Run database migrations as a separate, explicit step (a one-off ECS task or a pre-deploy job), and make them backward compatible. During a rolling deploy, old and new containers briefly coexist, so a migration that drops a column the old version still reads will cause an outage.

Deploying on AWS

Here is a deployment shape that scales comfortably from startup to serious volume:

  • Application Load Balancer in public subnets, fronted by AWS WAF with rate limiting on /webhooks/* and /v1/*.
  • ECS Fargate services in private subnets: payments-api (2+ tasks across AZs), payments-worker (scale on queue depth), and n8n (single main instance, or queue mode for heavier loads).
  • Amazon RDS for PostgreSQL, Multi-AZ, with automated backups and point-in-time recovery. Consider RDS Proxy if connection counts become a problem.
  • Amazon ElastiCache for Redis for the Socket.io adapter.
  • AWS Secrets Manager for gateway keys, injected into tasks at runtime, never baked into images.
  • A private, authenticated route to n8n: an internal ALB reachable only through your VPN or SSO. Do not expose the n8n editor to the public internet.
  • CloudWatch logs and alarms on 5xx rate, webhook queue age, and dead-letter count.

AWS deployment topology with ALB, WAF, ECS Fargate services, RDS PostgreSQL, ElastiCache Redis and private n8n

One small but vital detail for containers: handle SIGTERM. ECS sends it before stopping a task, and a graceful shutdown means in-flight webhook inserts and worker transactions finish cleanly.

// src/server.js (excerpt)
const shutdown = async () => {
  server.close();           // stop accepting new connections
  await pool.end();         // drain DB connections
  process.exit(0);
};
process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);

Observability and Security Essentials

A payment system you cannot see into is a liability. At minimum:

  • Structured logs (for example with pino) that include gateway, event_id, and payment_id on every line. Never log full card data, secrets, or raw signatures.
  • Metrics that matter: webhook receive-to-processed latency, count of retry and dead events, reconciliation drift per run, and per-gateway error rates.
  • Alerting on absence: if a gateway that normally sends hundreds of webhooks an hour sends none for 30 minutes, something is wrong, even if no error was logged.
  • Signature verification with constant-time comparison, plus timestamp tolerance where the gateway supplies one, to blunt replay attacks.
  • Least-privilege IAM roles per service, and rotated webhook secrets.
  • PCI scope reduction: use hosted checkout or client-side tokenization so raw card numbers never touch your servers.

Best Practices

  • Persist first, process later. The webhook handler's only job is to verify and store.
  • Make idempotency a database property. Unique constraints beat application checks every time.
  • Use integer minor units everywhere. Convert at the adapter boundary and nowhere else.
  • Model payment status as a forward-only state machine. Enforce it in SQL.
  • Keep gateway differences inside adapters. Business code should never import a vendor SDK.
  • Reconcile on a schedule. Webhooks can be delayed, dropped, or never sent. Polling the gateway for suspicious payments is your safety net.
  • Test with real PostgreSQL in CI. Concurrency bugs do not show up against mocks.
  • Version your internal event contract. Future you will want to add a fourth gateway without a rewrite.

Common Mistakes (and How to Avoid Them)

  1. Parsing the body before verifying the signature. express.json() re-serializes the body and breaks HMAC checks. Use express.raw() on webhook routes and mount it before any global JSON parser.
  2. Doing slow work inside the webhook request. Sending emails or calling other APIs before responding invites gateway timeouts and duplicate retries.
  3. Trusting event order. A failed event can arrive after succeeded. Guard transitions in SQL.
  4. Using floats for money. Cashfree's decimal amounts are the usual culprit. Parse as strings and convert to integers.
  5. Deduplicating with an in-memory set or a short-lived cache. It disappears on redeploy, and it does not work across containers.
  6. Failing over on ambiguous errors. A timeout is not a failure. Check before creating a second order.
  7. Making sockets the source of truth. Clients disconnect. Always allow a REST refetch.
  8. Giving n8n the keys to everything. Let it call your internal API. Keep secrets and money logic in versioned, tested code.
  9. Exposing the n8n editor publicly. It is an admin tool with powerful credentials. Keep it private.
  10. Skipping the dead-letter plan. Events that fail eight times need a human, a dashboard, and a replay button.

🚀 Pro Tips

  • Store the raw payload as JSONB, always. When a gateway changes a field name next year, you can replay old events through a fixed adapter.
  • Add a replay endpoint. A single internal route that resets a dead event to received has saved me hours during incidents.
  • Use pg_notify to wake the worker the instant an event lands, with a slow poll as a fallback. You get low latency without a broker.
  • Tag every outbound gateway request with your payments.id. Reconciliation becomes a trivial join instead of detective work.
  • Run a weekly "chaos webhook" test. Send duplicates, out-of-order events, and malformed signatures to staging and confirm nothing changes.
  • Cap n8n execution data retention. Payment payloads in execution logs can contain personal data, so prune aggressively and mask sensitive fields.
  • Start with PostgreSQL as your queue. Move to SQS or a dedicated broker only when metrics prove you need it.

📌 Key Takeaways

  • Accept webhooks fast, store them durably, and process them asynchronously through an inbox table.
  • Enforce idempotency with database constraints and a forward-only status state machine.
  • Hide Stripe, Razorpay, and Cashfree behind adapters so your core logic stays vendor-neutral.
  • Let n8n handle scheduling, retries, and alerting, while your API owns the money logic and secrets.
  • Use Socket.io for instant feedback, but always let clients recover truth from the REST API.
  • Ship with Docker, test against real PostgreSQL in CI, and deploy on ECS Fargate with Secrets Manager, RDS, and a private n8n.

Conclusion

Payments are one of the few areas in software where "it mostly works" is not acceptable, yet the usual path to production is a pile of vendor SDK calls and a cron job nobody wants to touch. The approach in this guide is different in one important way: it assumes things will go wrong. Webhooks will duplicate, arrive late, or never arrive. Gateways will have bad days. Deploys will happen mid-transaction.

By putting idempotency in PostgreSQL, isolating vendors behind adapters, separating accepting events from processing them, and handing recurring operational work to a visible tool like n8n, you get a system that fails safely and recovers on its own. Add real-time Socket.io updates and a disciplined CI/CD and AWS setup, and you have a billing backend that is calm under pressure.

Start small. Build the inbox table and one adapter, prove the pattern with a single gateway, then layer in the rest. The architecture stays the same whether you process ten payments a day or ten million.

References

Frequently asked questions

Why use more than one payment gateway in a SaaS product?

Three reasons: coverage, cost, and resilience. Razorpay and Cashfree give you local payment methods like UPI in India, Stripe covers cards and international markets, and a second gateway means an outage at one provider does not stop you from collecting revenue.

How do I make payment webhooks idempotent?

Store each verified event in a table with a unique constraint on the gateway name and the event ID, using INSERT ... ON CONFLICT DO NOTHING. Then make the state change itself safe to repeat by only allowing forward status transitions, for example pending to succeeded but never succeeded back to pending.

Should I process webhooks inside the HTTP request handler?

No. Verify the signature, write the raw event to the database, and return a 2xx immediately. Do the real work in a background worker. Gateways retry on slow or failed responses, and doing heavy work inline is the quickest way to create duplicate deliveries.

Is n8n production-ready for payment reconciliation?

Yes, when you run it self-hosted behind authentication, keep it in a private subnet, back it with PostgreSQL instead of SQLite, and let it call your own internal API rather than holding gateway secrets. It is excellent for scheduling, retries, and alerting. Keep the actual money logic in your codebase.

Where does Make.com fit if I already run n8n?

Make.com is a good fit for business-facing automations owned by finance or support teams, such as logging mismatches in a spreadsheet or opening a ticket. Engineering owns n8n for system-level jobs, and the two connect through simple webhooks.

How do I handle Cashfree and Razorpay amounts safely?

Always convert to integer minor units (paise or cents) at the adapter boundary. Razorpay already uses paise, Stripe uses the smallest currency unit, and Cashfree reports decimal values in major units, so parse those as strings and convert without floating-point math.

Discussion

All Articles
Payment Gateway IntegrationNode.jsExpress.jsPostgreSQLStripeRazorpayCashfreen8nWebhooksIdempotencySocket.ioAWSDockerCI/CDSaaS Billing

Written by

Niraj Kumar

Software Developer — building scalable systems for businesses.

Building this for real? TypeScript Full Stack Development — End-to-end TypeScript products — Node APIs, PostgreSQL/Prisma, auth, and typed frontends.