
Network Backbone Architect
- 24 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Design carrier and enterprise backbone networks: core/edge topology, OSPF/IS-IS/BGP routing policy, WAN/MPLS/SD-WAN, DCI, peering, and capacity planning.
About
Guides design of carrier- and enterprise-scale backbone networks covering topology, IGP/BGP routing policy, WAN/MPLS/SD-WAN, DCI, peering, QoS, and observability. A developer uses it when architecting core network routing, addressing plans, and resilience for multi-site or carrier networks.
- Covers OSPF, IS-IS, BGP design, route policy, ECMP, BFD, and EVPN/VXLAN spine-leaf
- Produces topology, routing design, WAN/SD-WAN, internet-edge, and capacity model outputs
Network Backbone Architect by the numbers
- 24 all-time installs (skills.sh)
- Ranked #804 of 1,039 Cloud & Infrastructure skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill network-backbone-architectAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 24 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Design carrier and enterprise backbone networks: core/edge topology, OSPF/IS-IS/BGP routing policy, WAN/MPLS/SD-WAN, DCI, peering, and capacity planning.
Files
Network Backbone Architect
When NOT to Use
- REST/GraphQL, ESB, or enterprise application integration design →
enterprise-integration-api-developer - Cloud Well-Architected, landing zone, and managed VPC/service selection →
cloud-architect,enterprise-cloud-architect - Day-two cloud resource configuration (subnets, LB, managed VPN to cloud) →
cloud-engineer - Cloud security guardrails, CSPM, and cloud IAM as primary deliverable →
cloud-security-engineer - Host firewall, endpoint, and corporate security control catalog →
information-security-engineer - OS patching, VM admin, and cloud instance operations →
cloud-system-administrator - Terraform modules, CI/CD, and K8s platform delivery without backbone routing design →
infrastructure-engineer - SLO programs, on-call, and incident response as the main task →
site-reliability-engineer - Cross-domain system ADRs unrelated to routing →
senior-system-architecture - Application throughput, caching, and horizontal scale without L3 design →
high-concurrency-scalability - Event bus and messaging topology only →
event-driven-architecture - OT/ICS segmentation, Purdue model, and plant protocols →
scada-ics-cyber-security-specialist
Related skills
| Need | Skill |
|---|---|
| Cloud reference architecture and hybrid connectivity hooks | cloud-architect |
| Enterprise cloud governance and multi-BU landing zones | enterprise-cloud-architect |
| Implement cloud networking and managed connectivity | cloud-engineer |
| Cloud network security controls and posture | cloud-security-engineer |
| IaC, physical DC build, and platform delivery | infrastructure-engineer |
| Reliability engineering, SLOs, and production incidents | site-reliability-engineer |
| Enterprise system architecture across domains | senior-system-architecture |
| Application-scale concurrency and load distribution | high-concurrency-scalability |
| Async messaging and event-driven integration | event-driven-architecture |
Core Workflows
1. Scope, constraints, and design principles
Clarify scale (sites, regions, carriers), traffic matrix, RTO/RPO for paths, and regulatory or sovereignty constraints.
See `references/network_backbone_architect_scope.md`.
2. Topology, hierarchy, and addressing
Define core/distribution/edge roles, summarization boundaries, and IP/VLAN/VRF plan.
See `references/topology_hierarchy_and_addressing.md`.
3. Routing protocols and policy
Select IGP (OSPF, IS-IS) and BGP design—peering, communities, path selection, and filtering.
See `references/routing_igp_bgp_and_policy.md`.
4. WAN, MPLS, and SD-WAN
Architect carrier services, underlay/overlay, hub-spoke vs full mesh, and SLA alignment.
See `references/wan_mpls_sdwan_and_carriers.md`.
5. DCI, peering, and internet edge
Design data center interconnect, IX/transit/peering, and anycast or multi-homing at the edge.
See `references/dci_peering_and_internet_edge.md`.
6. Resilience, QoS, capacity, and operations
Plan redundancy, BFD/FRR, backbone QoS, link sizing, change windows, and observability.
See `references/resilience_qos_capacity_operations.md`.
Outputs
- Backbone context — sites, traffic matrix, critical flows, and failure domains
- Logical topology — hierarchy, VRFs, summarization points, and DCI attachment
- Routing design — IGP areas/levels, BGP AS plan, policies, and community/tag semantics
- WAN/SD-WAN architecture — underlay, overlay, hub roles, and carrier map
- Internet edge brief — peering vs transit, IX placement, prefix origination, and filtering
- Resilience and QoS matrix — ECMP, BFD, FRR, DSCP classes, and maintenance domains
- Capacity model — link sizing assumptions, growth headroom, and trigger thresholds
- Observability plan — flow telemetry, SNMP/telemetry targets, and backbone dashboards
Principles
- Hierarchy before complexity — aggregate at core; keep edge policies simple
- Explicit failure domains — maintenance windows and blast radius per region or plane
- Policy at the edge — filter and tag at borders; keep core transit predictable
- Measure before oversizing — size links from busy-hour matrices and growth, not peak anecdotes
- Prefer L3 DCI — stretch L2 only with documented operational cost and risk
- Document one-way doors — ASN allocation, summarization boundaries, and peering contracts
DCI, peering, and internet edge
Table of contents
1. Data center interconnect goals 2. L3 DCI vs stretched L2 3. DCI transport options 4. EVPN multi-site considerations 5. Internet edge architecture 6. Peering, transit, and IX 7. Multi-homing and traffic engineering 8. Anycast and CDN adjacency 9. Edge security integration
Data center interconnect goals
DCI connects production DC pairs or active/active regions. Clarify application requirements before choosing technology:
| Requirement | Drives design toward… |
|---|---|
| Active/active L3 apps | BGP between DCs, summarization, ECMP |
| VM mobility / stretched cluster | L2 stretch (higher risk) |
| Sync replication low latency | Dedicated low-latency path, QoS |
| DR only | Asymmetric: backup DC cold/warm; simpler routing |
Document RPO/RTO for network path separately from storage replication.
L3 DCI vs stretched L2
| Approach | Pros | Cons |
|---|---|---|
| L3 DCI (routed) | Stable flooding domain; clear failure boundaries | App must tolerate routed hops |
| L2 stretch (VXLAN/OTV/VPLS) | Legacy apps needing L2 adjacency | Split brain, STP, broadcast storms |
| L3 with host routing (anycast GW) | Modern DC pattern | Requires orchestration discipline |
Default recommendation: L3 DCI with BGP between DC border leaves/routers; advertise summarized prefixes plus selective /32 for anycast services.
If L2 stretch required:
- Stretched VLAN only for defined VLAN list
- BUM handling and ARP suppression in EVPN
- Split-horizon and site-of-origin tagging
- Run GFDL or similar failure detection between sites
DCI transport options
| Transport | Characteristics |
|---|---|
| Dark fiber / DWDM | Lowest latency, customer-owned mux; capex heavy |
| Carrier Ethernet EPL | Fixed latency, point-to-point |
| MPLS pseudowire / EVPL | Flexible; watch oversubscription |
| IPsec/GRE over internet | Cost-effective; variable latency; encrypt |
| Wave service | Multiple lambdas; long lead times |
Diversity: dual paths with diverse physical routes (separate conduits, carriers, entrances).
Capacity: size for replication peak + headroom (see resilience reference); monitor utilization >70% as upgrade trigger.
EVPN multi-site considerations
When extending EVPN/VXLAN across DCs:
| Topic | Guidance |
|---|---|
| Control plane | eBGP EVPN between DC border leaves or route reflectors |
| Route types | Type-2 (MAC/IP), type-3 (IMET), type-5 (IP prefix) interop |
| BUM | Ingress replication vs multicast underlay |
| Multi-homing | ESIs for dual-attached hosts; avoid cross-site vMotion without design |
| Failure | Site isolation — withdraw type-3 routes on partition |
Coordinate with DC fabric team; backbone architect owns inter-DC BGP policy and WAN attachment, not server vSwitch config.
Internet edge architecture
Typical internet edge layers:
[ Peering / Transit routers ] <-- eBGP to carriers and IX
|
[ Edge firewalls / scrubbing ] <-- optional inline or divert
|
[ Core / DC border ] <-- iBGP or static to internalFunctions:
- Originate company prefixes with consistent AS_PATH
- Default route propagation to internal (full or partial) per policy
- DDoS scrubbing diversion (BGP RTBH or flowspec)
- DNS anycast and public service anycast if applicable
Place edge in DMZ VRF with controlled leaks to corp VRF.
Peering, transit, and IX
| Relationship | Description | When to use |
|---|---|---|
| Transit | Pay provider for full internet table or default | Baseline connectivity, backup |
| Peering (PNI) | Private bilateral with another network | High traffic ratio to one AS |
| IX peering | Shared L2 fabric at exchange point | Many peers, cost-efficient PNI |
| Route server | IX-facilitated BGP on shared fabric | Simplify multilateral peering |
IX checklist:
- Cross-connect or reseller port capacity
- LOA for MAC if required
- Prefix filters inbound/outbound per peer
- Maximum prefixes and monitoring
- Bogon and RPKI on all eBGP sessions
Peering policy document: which ASNs to accept, traffic ratios, settlement-free criteria.
Multi-homing and traffic engineering
Dual transit providers:
- Split prefix announcements — same NLRI to both with AS-path or prepend tuning
- Use communities received from providers to influence inbound
- Local-pref for outbound preference
Inbound engineering limitations: you control outbound local-pref; inbound depends on remote policy — use multiple locations and anycast for strong inbound control.
Partial routes: some enterprises take default only from transit and peering routes from IX — reduces table size on edge.
Anycast and CDN adjacency
| Service | Network pattern |
|---|---|
| Internal anycast DNS | Same /32 from multiple sites; IGP metric or BGP MED |
| Public anycast | Global announcement from multiple POPs; coordinate RPKI |
| CDN | CNAME to CDN; optional private interconnect to CDN (not full internet path) |
Document withdrawal procedure on site failure to avoid blackholing anycast elsewhere.
Edge security integration
Coordinate architecture (implementation may involve security engineering):
| Control | Placement |
|---|---|
| Stateful firewall | North-south at edge; avoid hairpin through distant DC |
| DDoS scrubbing | BGP divert to scrubbing center |
| RTBH | Trigger community at edge on attack destination |
| TLS inspection | Often at app layer; avoid breaking latency-sensitive paths |
Do not duplicate full GRC control catalog — align with cloud-security-engineer for cloud egress and information-security-engineer for corporate policy.
Logging: flow records (NetFlow/IPFIX) and BGP update logs at edge for peering disputes and incident response.
Network backbone architect scope
Table of contents
1. Role and boundaries 2. Stakeholders and inputs 3. Traffic matrix and critical flows 4. Design constraints 5. Deliverable checklist 6. Review gates 7. Handoffs to peer skills
Role and boundaries
The network backbone architect owns carrier- and enterprise-scale routed infrastructure between sites, data centers, and the public internet. Scope includes:
- Logical and physical hierarchy — core, distribution, aggregation, edge, and WAN attachment roles
- Routing and addressing — IGP domains, BGP policy, summarization, and multi-VRF/multi-tenant separation at the network layer
- WAN and carrier services — MPLS L3VPN, private lines, SD-WAN overlays, and hybrid underlay design
- DCI and internet edge — stretched fabrics only when justified; peering, transit, IX, and anycast placement
- Resilience and operations — ECMP, BFD, fast reroute, maintenance domains, QoS at backbone scope, capacity planning, and flow/SNMP/telemetry observability
Out of scope (defer to peer skills):
| Topic | Skill |
|---|---|
| Application APIs, ESB, message contracts | enterprise-integration-api-developer |
| Cloud landing zone, Well-Architected, service selection | cloud-architect, enterprise-cloud-architect |
| Cloud resource implementation | cloud-engineer |
| Endpoint and host security controls | information-security-engineer |
| OS/instance administration | cloud-system-administrator |
| Terraform/K8s delivery without routing design | infrastructure-engineer |
| SRE on-call and error budgets | site-reliability-engineer |
| OT/ICS and plant networks | scada-ics-cyber-security-specialist |
Stakeholders and inputs
Gather before locking topology:
| Stakeholder | Typical inputs |
|---|---|
| Application / platform teams | East-west vs north-south ratios, latency budgets, multicast needs |
| Security architecture | segmentation model, inspection points, DDoS and filtering requirements |
| Data center / facilities | rack power, cross-connect limits, diverse path availability |
| Carriers / IX | SLA, MTU, BGP session limits, community support, LOA timelines |
| Operations / NOC | change windows, tooling (SNMP, gNMI, NetFlow), escalation paths |
| Compliance | data residency, lawful intercept hooks, audit logging for config changes |
Minimum discovery artifacts:
- Site inventory with criticality tier (Tier 0 backbone vs Tier 2 branch)
- Current ASN, prefix holdings, and IRR/RPKI posture
- Existing IGP (if any) and pain points (slow convergence, area sprawl)
- RTO/RPO per site pair for network path loss (not application DR alone)
Traffic matrix and critical flows
Build a traffic matrix (source × destination × protocol × peak/average Mbps × packet size profile):
| Flow class | Examples | Design implications |
|---|---|---|
| Inter-site business | ERP, file shares, VoIP | Latency-sensitive; prefer stable IGP metrics |
| DCI replication | storage sync, DB log shipping | Loss/latency sensitive; may need dedicated DCI or QoS |
| Internet egress | SaaS, updates, guest Wi-Fi | Asymmetric; policy at edge; DDoS scrubbing |
| Cloud on-ramps | VPN/Direct Connect/ExpressRoute | Overlap with cloud-architect; document handoff points |
| Management | SNMP, SSH, NTP, syslog | Out-of-band or dedicated VRF; never compete with user QoS |
Label critical flows that must survive single-link or single-node failure within documented convergence time (e.g., sub-50 ms with BFD + FRR where required).
Design constraints
Document non-negotiables early:
- MTU end-to-end — jumbo on DCI vs 1500 on internet; TCP MSS clamping plan
- Address space — RFC1918 vs provider-independent; IPv6 dual-stack policy
- Vendor mix — single-vendor backbone vs multi-vendor with standardized features only
- Automation — NETCONF/gNMI, Git-backed config, and rollback expectations
- Regulatory — encryption in transit (MACsec/IPsec on WAN), geo restrictions on traffic paths
- Scale limits — max prefixes per peer, max ECMP width, max areas/levels in IGP
Deliverable checklist
| Deliverable | Contents |
|---|---|
| Context diagram | Sites, carriers, cloud on-ramps, security inspection |
| Logical topology | Layers, VRFs, summarization boundaries |
| IP plan | Allocations per site/VRF, loopbacks, link nets, anycast pools |
| Routing design | IGP type/areas, BGP AS plan, policies, communities |
| WAN architecture | Underlay, overlay, hub roles, carrier map |
| Internet edge | Peering/transit, IX list, prefix origination, filtering |
| Resilience matrix | Node/link SPOFs, BFD, FRR, maintenance domains |
| QoS design | DSCP classes, queue mapping on backbone hops |
| Capacity model | Link sizes, growth, upgrade triggers |
| Observability | Telemetry targets, flow export, alert thresholds |
| Migration waves | cutover order, rollback, parallel run duration |
Review gates
Run structured reviews before build:
1. Architecture review — hierarchy, summarization, and failure domains 2. Security review — filtering, RPKI, management plane isolation 3. Operations review — change windows, monitoring, runbooks for peer loss 4. Carrier review — LOA, BGP parameters, SLA alignment with design assumptions
Capture decisions in short ADRs: problem, options, decision, consequences, and revisit triggers.
Handoffs to peer skills
| When design touches… | Hand off to… |
|---|---|
| Account structure, shared services, hybrid cloud reference | cloud-architect, enterprise-cloud-architect |
| Implementing VPC, TGW, Cloud VPN, or managed LB | cloud-engineer |
| Terraform modules for network-adjacent infra | infrastructure-engineer |
| Production SLOs and incident playbooks | site-reliability-engineer |
| End-to-end system integration beyond L3 | senior-system-architecture |
Keep backbone documents routing-centric; cloud diagrams should reference backbone attachment points (peer IPs, VLANs, BGP neighbors) without duplicating full cloud service catalogs.
Resilience, QoS, capacity, and operations
Table of contents
1. Redundancy models 2. ECMP and hash polarization 3. BFD and fast detection 4. Fast reroute and FRR 5. Backbone QoS and DSCP 6. Capacity planning and link sizing 7. Maintenance domains and change windows 8. Observability architecture 9. Operational runbooks
Redundancy models
| Model | Description | Use case |
|---|---|---|
| 1+1 | Active/standby path or device | Legacy serial links, firewalls in A/S |
| 1:1 | Dedicated backup | Cost-sensitive remote sites |
| N+1 | Shared spare capacity | Core link groups |
| N+N (ECMP) | Active/active parallel paths | Modern backbone default |
Device redundancy:
- Dual supervisors on chassis where available
- vRR/NSR for control-plane restart
- Geographic redundancy — second DC or POP, not only second line card
Document SPOFs explicitly when cost prevents true diversity (single carrier entrance, single IX port).
ECMP and hash polarization
Equal-cost multipath load-shares flows across parallel links:
| Topic | Guidance |
|---|---|
| Hash fields | L3/L4 5-tuple typical; avoid polarizing on single field |
| Unequal cost | UCMP where supported for diverse bandwidth links |
| Polarization | Add link numbers or LAG member diversity to vary hash |
| Flow stickiness | Long flows stay on one path — size each member for peak single-flow + average |
LAG/MC-LAG: prefer L3 ECMP across routers over stretched L2 bundles when possible.
Verify symmetric routing with stateful firewalls in path.
BFD and fast detection
Bidirectional Forwarding Detection sub-second failure detection for:
- IGP (OSPF, IS-IS)
- BGP (especially eBGP on internet and MPLS CE)
- MPLS LSP (platform-dependent)
- LAG member links
| Parameter | Aggressive | Conservative |
|---|---|---|
| Tx/Rx interval | 50–300 ms | 500 ms–1 s |
| Multiplier | 3 | 3–5 |
Match carrier BFD support on MPLS handoffs. Disable aggressive BFD on high-latency satellite unless tested.
Fast reroute and FRR
| Mechanism | Layer | Benefit |
|---|---|---|
| LFA / RLFA | IGP | Local repair before global reconvergence |
| TI-LFA | IGP + SR | Broader repair coverage |
| BGP PIC | BGP | Precomputed backup next-hop |
| FRR for MPLS | MPLS | Fast switch to backup LSP |
Design steps:
1. Enable IGP fast convergence features platform supports 2. Verify micro-loop avoidance settings 3. Lab-test single fiber cut and node isolation
Document expected traffic shift during FRR (some micro-bursts acceptable).
Backbone QoS and DSCP
Define end-to-end DSCP model (example enterprise classes):
| Class name | DSCP (example) | Traffic |
|---|---|---|
| Network control | CS6 (48) | Routing protocols, BFD |
| Voice | EF (46) | RTP telephony |
| Video | AF41 (34) | Interactive video |
| Business critical | AF31 (26) | Tier-1 apps |
| Default | BE (0) | Standard |
| Scavenger | CS1 (8) | Backup, bulk |
Backbone queuing:
- Strict priority for voice (policed EF)
- WFQ/CBWFQ for assured and default
- Shape at WAN edge to carrier CIR
WAN carrier mapping: document DSCP → MPLS EXP or COS per carrier contract.
Avoid re-marking at every hop; trust boundaries at edge only unless mitigating trust issues.
Application QoS without network participation → coordinate with high-concurrency-scalability for app-level throttling; network still defines classes for real-time traffic.
Capacity planning and link sizing
Traffic engineering inputs
- Busy hour utilization per link (95th percentile)
- Growth rate (historical 6–12 months + business forecast)
- Burst factor for elephant flows (storage, backup windows)
Sizing formula (simplified)
required_bandwidth = peak_measured × (1 + growth%) × headroom_factorTypical headroom_factor: 1.3–1.5 for backbone; higher for internet unpredictable bursts.
Triggers
| Metric | Action |
|---|---|
| >70% util 5-min BH | Plan upgrade |
| >85% sustained | Emergency upgrade / traffic shift |
| Queue drops on PQ/WFQ | Review QoS or capacity |
| RTT/loss SLA miss | Path diversity or carrier escalation |
Maintain capacity dashboard per critical path; review quarterly.
Maintenance domains and change windows
Maintenance domain: set of nodes/links upgraded together with controlled traffic drain.
| Technique | Use |
|---|---|
| max-metric router-lsa / overload bit | Drain IGP before reboot |
| BGP graceful shutdown | Withdraw with GSHUT community |
| AS-path prepend outbound | Reduce inbound during window |
| Maintenance BGP community | Signal peers (if supported) |
Change classification:
| Class | Example | Window |
|---|---|---|
| Standard | ACL tune, description | Business hours with rollback |
| Major | IGP area redesign, RR move | Planned outage window |
| Emergency | Security filter for attack | Anytime with approval |
Document rollback (config snapshot, parallel run duration) for every backbone change.
Observability architecture
Architecture-level telemetry (implementation with NOC/tools):
| Source | Data | Use |
|---|---|---|
| NetFlow / IPFIX / sFlow | Per-flow stats | Capacity, anomaly, security |
| SNMP / gNMI telemetry | Interface counters, queues, CPU | Threshold alerts |
| BGP monitoring | Update rate, prefix count | Leak detection |
| Synthetic probes | TWAMP, ICMP, HTTP | Path SLA |
| Syslog / audit | Config changes | Compliance |
Key alerts (examples):
- eBGP session down > 1 min
- Prefix count deviation >10% from baseline
- Interface errors/CRC increment
- BFD session flap rate
- QoS drop counters on PQ
Correlate with site-reliability-engineer SLIs where application paths cross backbone.
Operational runbooks
Minimum runbook set:
| Scenario | Actions |
|---|---|
| Transit provider outage | Shift to alternate provider; verify prefix origination |
| IX port failure | Fail to backup cross-connect or transit |
| DCI link loss | Verify storage replication pause; IGP/BGP convergence |
| BGP prefix leak internal | Filter at RR; isolate source VRF |
| DDoS on public service | RTBH or scrubbing center diversion |
| Planned core upgrade | Drain via max-metric; verify ECMP shift |
Include escalation to carrier with circuit ID, LOA, and last-good config diff.
Store topology and policy source of truth in Git; tie to change tickets.
Review runbooks after post-incident review; feed architecture updates when recurring failures indicate design debt.
Routing IGP, BGP, and policy
Table of contents
1. Protocol selection 2. OSPF design 3. IS-IS design 4. IGP tuning and stability 5. BGP foundations 6. BGP policy toolkit 7. Route reflectors and confederations 8. Filtering and security 9. Convergence targets
Protocol selection
| Factor | Favor OSPF | Favor IS-IS |
|---|---|---|
| Team familiarity | Common in enterprise | Common in SP/large DC |
| TLV extensibility | Adequate | Strong for SR, flex-algo (if used) |
| Multi-area/level design | Areas 0 + non-backbone | Single protocol, wide metrics |
| IPv6 | OSPFv3 parallel or dual-stack | Often single protocol for v4/v6 |
| Vendor DC fabric | Less common as underlay IGP | Common underlay for spine-leaf |
Default guidance:
- Enterprise multi-site with mixed vendors: OSPF or IS-IS both viable; pick one per domain and avoid redistribution mess
- MPLS L3VPN underlay: often IS-IS or OSPF in provider core; customer CE uses static or BGP
- Internet edge: always BGP; never run IGP with external parties
OSPF design
Area design
- Area 0 connects all ABRs; keep area 0 small and stable
- Stub / totally stubby / NSSA at remote sites to shrink LSDB
- One area per site or per region depending on router count (rule of thumb: avoid >50 routers per area without modeling)
Types and metrics
- Use point-to-point on backbone links (no DR election on /31)
- Reference bandwidth aligned with fastest link in area
- BFD on all OSPF adjacencies on backbone (see resilience reference)
Redistribution
- Minimize static ↔ OSPF redistribution; use route tags and distribute-lists
- BGP → OSPF only at controlled borders with explicit prefix lists
IS-IS design
Levels
| Level | Role |
|---|---|
| L2 backbone | Core mesh or hub; carries summarized regional routes |
| L1 access | Site internal; default route to L1/L2 router |
| L1/L2 router | Site border; aggregates L1 into L2 |
NET addressing
- Consistent System ID and area plan; document in IPAM adjunct
- Wide metrics for 10G+ links; consider metric-style compatible with TE if Segment Routing deployed
Multi-topology
- Use only when IPv4/IPv6 topologies must diverge; otherwise single topology simplifies operations
IGP tuning and stability
| Knob | Purpose |
|---|---|
| Hello/dead intervals | Faster detection vs stability on lossy links |
| LSA/LSP pacing | Protect CPU during flaps |
| Max-metric router-lsa | Maintenance — drain traffic before change |
| Prefix suppression | Hide transit link prefixes if platform supports |
Graceful restart (NSF/NSR): document where enabled; verify helper mode on peers during upgrades.
Avoid IGP in the WAN cloud unless you own both ends; prefer BGP over MPLS or SD-WAN overlay.
BGP foundations
ASN plan
| ASN type | Use |
|---|---|
| Private (64512–65534, 4200000000–4294967294) | Internal iBGP |
| Public | Internet edge, IX peering, some DCI |
iBGP full mesh does not scale — use route reflectors (RR) or confederations.
Session types
| Session | Typical use |
|---|---|
| iBGP | Same AS; next-hop unchanged on RR with proper policy |
| eBGP | Internet, transit, IX, carrier MPLS VPN (if applicable) |
| BGP labeled unicast | MPLS VPN or SR transport (carrier-dependent) |
Address families
- IPv4 unicast — default everywhere
- IPv6 unicast — parallel policy; separate peer policies if needed
- VPNv4/VPNv6 — MPLS L3VPN to CE or between PEs
- EVPN — DCI or DC fabric (type-2/3/5 routes)
BGP policy toolkit
Express policy with route-maps, prefix-lists, AS-path filters, communities:
| Mechanism | Example use |
|---|---|
| Local preference | Prefer one transit or one DCI path |
| MED | Influence inbound from dual-homed site (same AS only) |
| AS-path prepend | Deprioritize outbound path (limited effect) |
| Communities | Signal regional preference, blackhole, no-export |
| ORF / prefix limit | Cap prefixes accepted from peer |
Community cookbook (document organization-specific values):
| Community | Meaning |
|---|---|
| 65000:100 | Prepend once toward internet |
| 65000:666 | Blackhole / RTBH trigger at edge |
| 65000:90 | Do not export to IX peers |
Align inbound policy: max-prefix, bogon filter, RPKI ROV where supported.
Route reflectors and confederations
Route reflector cluster:
- Place RR at distribution or dedicated RR pair
- Cluster ID per RR pair; clients only peer to RR (or local cluster)
- Enable next-hop-self on RR only where CE/edge needs it
Confederation: split AS into sub-AS for very large networks; use when RR alone insufficient.
RR on WAN: avoid reflecting across high-latency links without modeling; regional RR pairs common.
Filtering and security
| Location | Minimum controls |
|---|---|
| Internet eBGP | Bogon + martian, max-prefix, RPKI invalid drop (if available) |
| IX peering | Prefix limits per LOA; strict inbound filter from peer |
| iBGP | TTL security (GTSM) on loopback sessions where supported |
| VPN PE-CE | CE cannot originate full table; default or limited prefixes |
RPKI: document ROA coverage for originated prefixes; monitor INVALID state.
Management plane: BGP sessions from loopbacks in MGMT VRF where possible.
Convergence targets
Document expected behavior:
| Event | Target (example) | Mechanisms |
|---|---|---|
| Link failure | < 1 s | BFD + IGP fast hello; BGP PIC |
| Node failure | < 3 s | ECMP removal; FRR (LFA/RLFA) |
| Peer loss (internet) | < 30 s | BGP hold timer; alternate path |
Validate in lab or maintenance window with controlled flap tests.
Pair operational runbooks with site-reliability-engineer for application-visible SLOs—not network-only metrics alone.
Topology hierarchy and addressing
Table of contents
1. Hierarchical reference model 2. Core, distribution, and edge roles 3. Spine-leaf and EVPN context 4. VRF and multi-tenant segmentation 5. Addressing plan 6. Summarization strategy 7. Loopbacks and anycast 8. Documentation conventions
Hierarchical reference model
Classic three-tier campus/enterprise hierarchy maps to WAN/backbone as follows:
[ Internet / IX / Transit ]
|
+---------+---------+
| Internet Edge |
+---------+---------+
|
+---------+---------+
| WAN / MPLS Core | <- often provider or owned core
+---------+---------+
/ | \
+------------+ | +------------+
| DC / Site A | Site B |
| Distribution | Distrib. |
| | | |
| Access/Edge | Access |
+-------------------+--------------+Design intent:
- Core — high-speed transit only; minimal policy; maximum summarization inward
- Distribution — policy aggregation, route reflection, firewall insertion, WAN handoff
- Edge — access, local default, local summarization toward distribution
Avoid flat full-mesh at scale; use hierarchy to bound IGP LSDB and policy complexity.
Core, distribution, and edge roles
| Layer | Functions | Typical devices | Routing notes |
|---|---|---|---|
| Core | Non-stop forwarding between major hubs | chassis routers, high-density fabrics | IGP on /31 or /30 only; no end-host routes |
| Distribution | Site aggregation, DC border, RR placement | modular switches/routers | Summarize site prefixes; BGP to WAN/core |
| Edge | Access VLANs, WAN CPE, branch | switches, SD-WAN edge, CPE | Default route or partial routes; stub IGP |
Dual-homing rules:
- Every distribution node dual-attaches to diverse core paths where physically possible
- Edge uses two upstreams (distribution pair or active/active SD-WAN) with consistent metrics
- Document primary/secondary only when using active/standby; prefer ECMP where protocols allow
Spine-leaf and EVPN context
For data center and large campus, spine-leaf with EVPN/VXLAN often replaces classic three-tier inside the DC. Backbone architect scope:
| Inside DC (spine-leaf) | At DC border (backbone attachment) |
|---|---|
| VXLAN VNI, EVPN type-2/3/5 | BGP EVPN to core or route reflectors |
| Anycast gateway on leaves | Summarize tenant prefixes at border leaf pair |
| Clos ECMP east-west | L3 handoff to WAN (no unnecessary L2 stretch) |
DCI caution: extending L2 VXLAN between DCs requires explicit design (multi-site EVPN, stretched VLAN). Prefer L3 DCI (host routes or summarized prefixes over BGP) unless application mandates L2 adjacency.
Cross-link: application scale patterns → high-concurrency-scalability (not a substitute for L3 design).
VRF and multi-tenant segmentation
Use VRFs (or equivalent logical routers) to separate:
| VRF | Typical contents |
|---|---|
| GLOBAL / default | Internet routing, shared services |
| CORP | desktops, internal apps |
| DMZ | public-facing app tiers |
| MGMT | OOB, jump hosts, device management |
| GUEST | captive portal, internet-only |
| PARTNER | extranet B2B |
Leakage control:
- Route targets / import-export policies explicit in BGP VPN
- No accidental full table leak between VRFs at RR
- Management VRF reachable only via controlled jump paths
Cloud hybrid: map VRF intent to cloud VPC/VNet segmentation in joint workshops with cloud-architect.
Addressing plan
IPv4 private space
Allocate by region → site → function with reserved growth:
| Block tier | Example use |
|---|---|
| /16 per region | All sites in geography |
| /20 per site | Loopbacks, links, summaries |
| /24 per function | MGMT, prod, DMZ within site |
Use consistent loopback addressing (/32 per device) for BGP router-id and telemetry source.
IPv6
If dual-stack:
- ULA or GUA per organizational policy
- Parallel summarization at same hierarchy points as v4
- Verify PMTUD and extension header handling on WAN
Link addressing
- Point-to-point
/31(v4) or/127(v6) on backbone links - Document interface descriptions schema:
CORE1-DIST2_lag12_10G
Summarization strategy
Summarize as close to the source as operationally safe:
| Location | What to summarize |
|---|---|
| Edge | Access subnets → single aggregate per building |
| Distribution | Site aggregate toward core/WAN |
| Core | Regional supernets toward other regions |
Stability vs agility tradeoff:
- More specific summary → faster blackhole risk if summary anchor fails
- Less specific → larger tables and suboptimal hot-potato
Use null route or discard on summary anchor with more-specific leak only for critical /32 exceptions (document each exception).
Loopbacks and anycast
| Type | Purpose |
|---|---|
| Physical loopback / system IP | Stable BGP RID, SNMP source, SSH target |
| Service anycast | DNS, load balancer VIP, NTP in multiple sites |
Anycast on backbone:
- Same prefix announced from multiple sites with consistent community or MED policy
- Tie health to IGP or BGP withdraw on failure
- Coordinate with
dci_peering_and_internet_edge.mdfor internet anycast vs internal anycast
Documentation conventions
Maintain living artifacts:
- IPAM spreadsheet or IPAM tool export — authoritative allocations
- VRF matrix — VRF × site × import/export targets
- Cable / port plan — only where needed for backbone handoffs (full physical design may sit with
infrastructure-engineer) - Naming — hostname, interface, and BGP neighbor naming standards
Reconcile quarterly: orphaned subnets, summary drift, and prefix bloat toward internet edge.
WAN, MPLS, SD-WAN, and carriers
Table of contents
1. WAN architecture patterns 2. Private line and Ethernet services 3. MPLS L3VPN 4. SD-WAN overlay 5. Hybrid underlay and overlay 6. Carrier selection and SLA 7. MTU and fragmentation 8. Migration and coexistence
WAN architecture patterns
| Pattern | Topology | Best for | Risks |
|---|---|---|---|
| Hub-and-spoke | Branches → regional hub → core | Cost control, centralized inspection | Hub SPOF, latency hairpin |
| Partial mesh | Meshed hubs, spokes to hub | Balance resiliency and cost | Policy complexity |
| Full mesh | All sites meshed | Lowest latency between all pairs | Cost, table size, operations |
| Regional hub | Continent hubs with inter-hub links | Global enterprises | Inter-hub capacity planning |
Choose based on traffic matrix, not diagram aesthetics. Model hub failure and carrier cut explicitly.
Private line and Ethernet services
Dedicated circuits (TDM legacy, Ethernet EPL/EVPL):
| Attribute | Design note |
|---|---|
| Bandwidth | Committed vs oversubscribed (especially EVPL) |
| Diversity | Diverse POP paths into building; diverse CPE |
| Handoff | Copper vs fiber; single-mode vs multimode |
| SLA | Availability, latency, jitter for voice/video |
| Lead time | Long — plan architecture early |
Point-to-point Ethernet between DCs often forms DCI underlay before MPLS or IPsec overlay.
Document demarcation and who owns CPE vs NID.
MPLS L3VPN
Carrier provides VPNv4/VPNv6 with customer VRFs mapped to route targets (RT).
| Role | Device | Function |
|---|---|---|
| PE | Provider edge | Terminates VPN; imposes labels |
| P | Provider core | Label switch only |
| CE | Customer edge | BGP or static to PE |
Design checklist:
- Unique RD per VRF per site (or per service) per carrier guidance
- Import/export RT matrix documented
- BGP CE–PE: default route vs full routes vs selective prefixes
- QoS marking honored in carrier COS (verify DSCP→EXP mapping)
- Multicast rarely supported — confirm before design depends on it
Inter-provider VPN (option B/C): only when multi-carrier strategy requires; high operational cost.
SD-WAN overlay
SD-WAN builds encrypted overlay (often IPsec/GRE) across any underlay (broadband, LTE, MPLS).
| Component | Responsibility |
|---|---|
| Orchestrator | Central policy, zero-touch provisioning |
| Edge appliance / vCPE | Local breakout, path selection, DPI (vendor-dependent) |
| Underlay | MPLS, DIA, LTE — diverse paths recommended |
Path selection policies:
- Application-aware routing based on SLA (latency, loss, jitter)
- Local internet breakout for SaaS vs backhaul to hub
- FEC / packet duplication for lossy links (bandwidth cost)
Architectural decisions:
| Question | Options |
|---|---|
| Hub vs mesh overlay | Meshed edges for DC-like sites; hub for branches |
| Control plane | BGP to LAN, static, or dynamic to DC |
| Integration with MPLS | Hybrid: MPLS primary, DIA backup underlay |
| Security | ZTNA/SWG integration at edge — coordinate with security architecture |
Avoid double NAT and asymmetric routing when mixing breakout and hub inspection.
Hybrid underlay and overlay
Common enterprise pattern:
[ Branch SD-WAN ] ---- LTE / DIA ----+
| |
+---- MPLS VPN (COS gold) -----+----> [ Regional Hub / DC ]Design rules:
- Underlay diversity — at least two independent paths per site (different carriers or media)
- Consistent addressing — overlay uses same corporate VRF semantics as MPLS
- Routing — redistribute sparingly between overlay and MPLS; prefer single control point at hub
- QoS — end-to-end class only if underlay honors DSCP; else application SLA drives path pick only
Carrier selection and SLA
Evaluate carriers on:
| Criterion | Notes |
|---|---|
| Footprint | On-net sites vs expensive last-mile build |
| SLA credits | Availability %, MTTR, latency guarantees |
| BGP features | Communities, BFD, max-prefix, dual-stack |
| Support | 24×7 NOC, escalation, maintenance notification lead time |
| Security | DDoS scrubbing offering, RTBH support |
| Financial | contract term, burstable vs flat rate |
Maintain carrier scorecard and secondary carrier for critical sites.
MTU and fragmentation
| Layer | Typical MTU | Action |
|---|---|---|
| Internet DIA | 1500 | TCP MSS clamp at edge if tunneling |
| MPLS | 1508–9192 | Confirm end-to-end; jumbo only if full path supports |
| SD-WAN IPsec | 1400–1450 effective | Lower LAN MTU or clamp MSS |
| Overlay inside overlay | Lowest MTU wins | Model in lab |
Document DF bit policy and PMTUD black hole detection.
Migration and coexistence
MPLS to SD-WAN migration waves:
1. Pilot sites with parallel run (MPLS + overlay) 2. Shift default route to overlay; keep MPLS for critical classes 3. Decommission MPLS tail after soak period
Rollback: retain MPLS VLAN/circuit until overlay stable 30–90 days.
Coordinate cutovers with maintenance domains in resilience reference; notify application owners for TCP long-flow resets.
Cloud WAN: hand off cloud-side VPN/Direct Connect attachment design to cloud-engineer with backbone BGP parameters documented.