Advanced Golang: building a real payment processor
Hey everyone!
Over the past few months I put together a course about building a payment processor in Go. Not a simplified example: four microservices, architecture decisions documented as ADRs, and a full deploy on GKE.
The course is on Udemy: GoLang Avançado: Microsserviços, gRPC, Kafka e IA Antifraude.
Here is what is inside.
The complete system
Five components, each with a clear responsibility:
Native fraud detection in Go
The fraud service runs in the synchronous authorization path with a budget of around 15ms. The usual setup is a Python sidecar or a dedicated inference server: one more network hop, one more runtime to operate, one more serialization per call.
Gradient boosted trees (the standard model family for tabular fraud data) reduce at inference time to comparisons and additions. Nothing about them requires an ML runtime.
The course implements inference directly in Go. The model is exported as a JSON file of flat node arrays, validated at startup, and walked directly. The scorer allocates nothing on the hot path and sits behind a generic dynamic batcher. The benchmark came out at ~8 nanoseconds per score with zero allocations.
FSM with deterministic fallback
The fraud pipeline has five named states:
The key point about the fallback: if ml_score fails or exceeds its budget, the system does not stall the authorization. The handler marks the decision as degraded, routes it to manual review, and continues. Availability wins over model coverage, and the degradation is visible in the response instead of being silent.
Generic dynamic batcher
The tree scorer sits behind a batcher that accumulates requests over a short time window before processing them in bulk. The implementation uses generics:
1
2
3
4
5
6
7
8
9
10
// the batcher is parameterized on request and response types
type Batcher[Req, Resp any] struct {
window time.Duration
process func([]Req) []Resp
// ...
}
func (b *Batcher[Req, Resp]) Submit(ctx context.Context, req Req) (Resp, error) {
// groups concurrent calls and dispatches in batch
}
The same Batcher serves the fraud scorer, embedding calls, and any other hot path that needs batching, without duplicating the windowing logic.
MCP from scratch in Go
The investigator exposes three investigation tools:
| Tool | What it does |
|---|---|
get_account_history | payment history for the account |
geo_lookup | IP geolocation (ipapi.co or static) |
search_fraud_patterns | similarity search via pgvector/cosine |
The MCP server is implemented directly on top of encoding/json and stdio, with no SDK and no external dependencies. The full protocol (JSON-RPC 2.0, initialize handshake, tool listing, tool_call semantics) is course material.
The same tool registry serves two frontends: the stdio MCP for external clients (any MCP client can connect with go run ./cmd/investigator --mcp) and in-process calls for the investigation advisors.
By default, the investigator runs a rule-based advisor. When ANTHROPIC_API_KEY is set, the same tool registry is passed to a Claude agent, without changing anything in the registry itself.
Outbox: consistency without two-phase commit
Event publishing is an area where many systems have silent bugs. The problematic sequence:
1
2
3
4
5
1. BEGIN
2. INSERT payment
3. COMMIT
4. process dies here
5. publish to Kafka ← never happens
The outbox pattern fixes this by writing the event in the same transaction as the payment:
1
2
3
4
5
6
7
8
1. BEGIN
2. INSERT payment
3. INSERT outbox_events (same tx)
4. COMMIT
← if the process dies here, the relay recovers on the next poll
5. relay: SELECT unpublished FROM outbox_events
6. publish to Kafka
7. UPDATE outbox_events SET published=true
The relay polls every 200ms, publishes, and marks as published after the broker ACKs. At-least-once delivery, idempotent consumers keyed by payment_id.
Caches without Redis
Velocity counting and the semantic cache are kept in-process, without Redis. Two reasons from the ADR:
- A network round-trip on the authorization hot path costs more than the cache saves
- Redis goes down, authorization degrades
The implementation uses types with manual sharding:
1
2
3
4
5
6
7
8
9
10
// velocity: sharded TTL map for per-card attempt counting
type VelocityCounter[K comparable] struct {
shards [N]shard[K]
}
// semantic cache: rolling window of embedding vectors
type SemanticCache struct {
embeddings []embeddingEntry
mu sync.RWMutex
}
With multiple fraud replicas, each holds independent state. Load balancing via card fingerprint hash at the gRPC client softens this.
pgvector: similarity search in Postgres
The investigator searches historical similar cases to provide context for the investigation. The decision was to use pgvector in the same Postgres the ledger already uses, rather than spinning up a dedicated vector database.
1
2
3
4
5
-- find the 5 most similar cases by cosine distance
SELECT id, description, decision
FROM fraud_patterns
ORDER BY embedding <=> $1 -- cosine distance operator
LIMIT 5;
The index is IVFFlat for approximate nearest neighbor search. The PatternStore port isolates the implementation, so swapping to Qdrant means changing one adapter.
Observability
Every service exposes /metrics in Prometheus format. Grafana has a dashboard provisioned as code in the repository, and the same manifests run on docker-compose, kind, and the GCP cluster.
The operator console (localhost:3001) is a pre-built web interface showing the live payment feed and review queue. It connects to the gateway via SSE and degrades progressively: each panel shows a locked state with the course module name that unlocks it.
1
2
make dashboards # opens Grafana and Prometheus
make demo-feed # global stream of payments and reviews
Optional integrations
Everything runs offline by default. Integrations activate by environment variable:
Without any of these keys, the system runs complete with the rule-based advisor, static geo, and no external notifications.
GKE deployment
Terraform provisions two node pools on GKE: a standard pool and a dedicated pool for the fraud service, with taint and autoscaling from 1 to 8 replicas on ARM instances. The database uses Cloud SQL for PostgreSQL 16 with private IP and native pgvector.
1
2
3
4
make up # full local stack with docker-compose
# --- after covering the infra module ---
terraform init && terraform apply # GKE + Cloud SQL + Managed Kafka
kubectl apply -k deploy/k8s/overlays/gcp
ADRs
Every technical decision in the system has an ADR documenting context, decision, and consequences (positive and negative). There are 11 ADRs in the repository, covering everything from the choice of Postgres as the single database to the console as a walking skeleton.
What you will build
Advanced Go
Generics, goroutines, channels, sync.RWMutex, context, idiomatic errors, type-parameterized dynamic batcher.
Native inference
Decision tree in Go, ~8ns per score, zero allocations on the hot path. No Python sidecar.
Auditable FSM
Fraud pipeline modeled as an explicit FSM. Deterministic fallback, per-transition trace, loops structurally impossible.
Protocol Buffers and gRPC
Service contracts from scratch. HTTP, SSE, and gRPC in a single gateway. SSE for the operator console.
Kafka and Outbox
Consistency between Postgres and Kafka with outbox, idempotency by payment_id, double-entry accounting.
MCP from scratch
MCP server in JSON-RPC 2.0 over stdio. Claude agent or rule-based advisor — same tool registry.
pgvector
Similarity search in the same Postgres as the ledger. IVFFlat index, cosine distance, port to swap for Qdrant.
GKE + Terraform
Kubernetes deploy on GCP. Two node pools, Cloud SQL, Managed Kafka, Grafana provisioned as code.
There is a YouTube presentation video if you want to see the system before enrolling: YouTube presentation.
Course link: GoLang Avançado: Microsserviços, gRPC, Kafka e IA Antifraude.
See you in the next post!
