Post
Available in: Português

Advanced Golang: building a real payment processor

Hey everyone!

Over the past few months I put together a course about building a payment processor in Go. Not a simplified example: four microservices, architecture decisions documented as ADRs, and a full deploy on GKE.

The course is on Udemy: GoLang Avançado: Microsserviços, gRPC, Kafka e IA Antifraude.

Here is what is inside.


The complete system

Five components, each with a clear responsibility:

Payment Processor — overview client HTTP console :3001 gateway HTTP + SSE rate limiting gRPC client :8080 gRPC gRPC ledger double-entry idempotency Postgres 16 outbox → Kafka :9001 fraud decision tree ~8ns per score FSM + fallback semantic cache :9002 Kafka Redpanda investigator Claude agent MCP server pgvector RAG Slack + email Kafka consumer Postgres 16 + pgvector Prometheus + Grafana · GKE (Cloud SQL · Managed Kafka · Managed Prometheus)

Native fraud detection in Go

The fraud service runs in the synchronous authorization path with a budget of around 15ms. The usual setup is a Python sidecar or a dedicated inference server: one more network hop, one more runtime to operate, one more serialization per call.

Gradient boosted trees (the standard model family for tabular fraud data) reduce at inference time to comparisons and additions. Nothing about them requires an ML runtime.

The course implements inference directly in Go. The model is exported as a JSON file of flat node arrays, validated at startup, and walked directly. The scorer allocates nothing on the hot path and sits behind a generic dynamic batcher. The benchmark came out at ~8 nanoseconds per score with zero allocations.


FSM with deterministic fallback

The fraud pipeline has five named states:

Fraud pipeline — FSM states enrich geo + velocity fast_rules blocklist · caps ml_score tree ~8ns semantic_tie pgvector cache finalize approve / deny fallback: degraded → manual review every transition recorded with duration and error · loops structurally impossible

The key point about the fallback: if ml_score fails or exceeds its budget, the system does not stall the authorization. The handler marks the decision as degraded, routes it to manual review, and continues. Availability wins over model coverage, and the degradation is visible in the response instead of being silent.


Generic dynamic batcher

The tree scorer sits behind a batcher that accumulates requests over a short time window before processing them in bulk. The implementation uses generics:

1
2
3
4
5
6
7
8
9
10
// the batcher is parameterized on request and response types
type Batcher[Req, Resp any] struct {
    window  time.Duration
    process func([]Req) []Resp
    // ...
}

func (b *Batcher[Req, Resp]) Submit(ctx context.Context, req Req) (Resp, error) {
    // groups concurrent calls and dispatches in batch
}

The same Batcher serves the fraud scorer, embedding calls, and any other hot path that needs batching, without duplicating the windowing logic.


MCP from scratch in Go

The investigator exposes three investigation tools:

ToolWhat it does
get_account_historypayment history for the account
geo_lookupIP geolocation (ipapi.co or static)
search_fraud_patternssimilarity search via pgvector/cosine

The MCP server is implemented directly on top of encoding/json and stdio, with no SDK and no external dependencies. The full protocol (JSON-RPC 2.0, initialize handshake, tool listing, tool_call semantics) is course material.

The same tool registry serves two frontends: the stdio MCP for external clients (any MCP client can connect with go run ./cmd/investigator --mcp) and in-process calls for the investigation advisors.

By default, the investigator runs a rule-based advisor. When ANTHROPIC_API_KEY is set, the same tool registry is passed to a Claude agent, without changing anything in the registry itself.


Outbox: consistency without two-phase commit

Event publishing is an area where many systems have silent bugs. The problematic sequence:

1
2
3
4
5
1. BEGIN
2. INSERT payment
3. COMMIT
4. process dies here
5. publish to Kafka  ← never happens

The outbox pattern fixes this by writing the event in the same transaction as the payment:

1
2
3
4
5
6
7
8
1. BEGIN
2. INSERT payment
3. INSERT outbox_events  (same tx)
4. COMMIT
                         ← if the process dies here, the relay recovers on the next poll
5. relay: SELECT unpublished FROM outbox_events
6. publish to Kafka
7. UPDATE outbox_events SET published=true

The relay polls every 200ms, publishes, and marks as published after the broker ACKs. At-least-once delivery, idempotent consumers keyed by payment_id.


Caches without Redis

Velocity counting and the semantic cache are kept in-process, without Redis. Two reasons from the ADR:

  1. A network round-trip on the authorization hot path costs more than the cache saves
  2. Redis goes down, authorization degrades

The implementation uses types with manual sharding:

1
2
3
4
5
6
7
8
9
10
// velocity: sharded TTL map for per-card attempt counting
type VelocityCounter[K comparable] struct {
    shards [N]shard[K]
}

// semantic cache: rolling window of embedding vectors
type SemanticCache struct {
    embeddings []embeddingEntry
    mu         sync.RWMutex
}

With multiple fraud replicas, each holds independent state. Load balancing via card fingerprint hash at the gRPC client softens this.


pgvector: similarity search in Postgres

The investigator searches historical similar cases to provide context for the investigation. The decision was to use pgvector in the same Postgres the ledger already uses, rather than spinning up a dedicated vector database.

1
2
3
4
5
-- find the 5 most similar cases by cosine distance
SELECT id, description, decision
FROM fraud_patterns
ORDER BY embedding <=> $1  -- cosine distance operator
LIMIT 5;

The index is IVFFlat for approximate nearest neighbor search. The PatternStore port isolates the implementation, so swapping to Qdrant means changing one adapter.


Observability

Every service exposes /metrics in Prometheus format. Grafana has a dashboard provisioned as code in the repository, and the same manifests run on docker-compose, kind, and the GCP cluster.

The operator console (localhost:3001) is a pre-built web interface showing the live payment feed and review queue. It connects to the gateway via SSE and degrades progressively: each panel shows a locked state with the course module name that unlocks it.

1
2
make dashboards   # opens Grafana and Prometheus
make demo-feed    # global stream of payments and reviews

Optional integrations

Everything runs offline by default. Integrations activate by environment variable:

ANTHROPIC_API_KEY — Claude agent in the investigator
STRIPE_API_KEY — real charges in test mode
SLACK_WEBHOOK_URL — investigation reports to Slack
RESEND_API_KEY — investigation reports by email
GEO_PROVIDER=ipapi — real geolocation via ipapi.co

Without any of these keys, the system runs complete with the rule-based advisor, static geo, and no external notifications.


GKE deployment

Terraform provisions two node pools on GKE: a standard pool and a dedicated pool for the fraud service, with taint and autoscaling from 1 to 8 replicas on ARM instances. The database uses Cloud SQL for PostgreSQL 16 with private IP and native pgvector.

1
2
3
4
make up           # full local stack with docker-compose
# --- after covering the infra module ---
terraform init && terraform apply   # GKE + Cloud SQL + Managed Kafka
kubectl apply -k deploy/k8s/overlays/gcp

ADRs

Every technical decision in the system has an ADR documenting context, decision, and consequences (positive and negative). There are 11 ADRs in the repository, covering everything from the choice of Postgres as the single database to the console as a walking skeleton.


What you will build

concurrency

Advanced Go

Generics, goroutines, channels, sync.RWMutex, context, idiomatic errors, type-parameterized dynamic batcher.

fraud

Native inference

Decision tree in Go, ~8ns per score, zero allocations on the hot path. No Python sidecar.

architecture

Auditable FSM

Fraud pipeline modeled as an explicit FSM. Deterministic fallback, per-transition trace, loops structurally impossible.

protocol

Protocol Buffers and gRPC

Service contracts from scratch. HTTP, SSE, and gRPC in a single gateway. SSE for the operator console.

events

Kafka and Outbox

Consistency between Postgres and Kafka with outbox, idempotency by payment_id, double-entry accounting.

ai

MCP from scratch

MCP server in JSON-RPC 2.0 over stdio. Claude agent or rule-based advisor — same tool registry.

database

pgvector

Similarity search in the same Postgres as the ledger. IVFFlat index, cosine distance, port to swap for Qdrant.

infra

GKE + Terraform

Kubernetes deploy on GCP. Two node pools, Cloud SQL, Managed Kafka, Grafana provisioned as code.


There is a YouTube presentation video if you want to see the system before enrolling: YouTube presentation.

Course link: GoLang Avançado: Microsserviços, gRPC, Kafka e IA Antifraude.

See you in the next post!