Tactical DDD with rich aggregates and value objects¶
Context and problem statement¶
The rules that matter most in Incident Ops are small and easy to get subtly wrong: which status transitions are allowed, when an acknowledgement window opens and closes, when escalation stops, when an incident complies with its SLA, when a check that fires late must do nothing. The same rules are read by three services. Where do these rules live, and how is it guaranteed that they are not reimplemented in endpoints, handlers or triggers?
Decision drivers¶
- One implementation per rule inside each bounded context, testable without a database, a host or a network.
- Invalid states (level 4, an empty title, a naive timestamp, a negative window) cannot be represented.
- Endpoints, triggers and routers stay thin enough to review at a glance.
- The structure is checkable by a build step, not only by review.
Considered options¶
- Rich aggregates and value objects per bounded context, with domain events and use cases that only orchestrate
- Anemic entities with service classes holding the rules (transaction script)
- One shared domain library used by every service
Decision outcome¶
Chosen option: rich aggregates and value objects per bounded context.
- Incident Management (
services/api):Incidentis the aggregate root; it owns itsTimelineEntryentities and anSlaClockvalue object, enforces transitions and raises one domain event per behavior.ServiceandOnCallRotationare separate aggregates. Repositories exist only for roots. - Escalation (
services/functions):AcknowledgementWatchis the aggregate;EscalationLevel,AcknowledgementDeadlineandSeveritycarry the rules; decisions are returned asEscalationDecisionvalues. - Operational Analytics (
services/insights):IncidentRecordandIncidentHistoryanswer measurement questions; value objects (ReportingWindow,Percentage,SlaPolicy) reject invalid input; domain services hold computation across many records. - Use cases load, call one behavior and commit or act through ports. Architecture tests (NetArchTest, reflection) and import-linter contracts fail the build when a layer reaches the wrong way or a domain type gains a public setter.
Consequences¶
- Good: the rules are unit tested in milliseconds (109 domain tests in the API, decision tables for
AcknowledgementWatch.EvaluateandPagingDecision.Decide). - Good: handlers, triggers and routers have no branches about the domain, so review focuses on wiring and tests.
- Good: each context keeps its own language; the contract is the only shared artifact (context map).
- Bad: more types and mapping code: EF Core value converters and an owned type for
SlaClock, anti-corruption translators in Escalation and Analytics. - Bad: the SLA compliance rule exists in Incident Management and in Analytics. Both follow the contract definition and are pinned by tests in each module.
Pros and cons of the options¶
Anemic entities with services¶
- Good: fewer types; maps directly to tables.
- Bad: nothing stops a second service from changing state without the rule; invariants depend on discipline.
Shared domain library¶
- Good: one copy of the SLA rule.
- Bad: couples release cycles of .NET and Python consumers (Python cannot use it at all); every context inherits the vocabulary of the system of record.