Skip to main content

Operating Routed Networks

Routed Networks connect an entire Project VPC to approved external IPv4 destinations through platform-managed eBGP gateways. They are for corporate LANs, physical services, and other routed domains that are not the Internet.

This is deliberately different from a Datacenter VLAN:

Routed NetworkDatacenter VLAN
RelationshipL3 route for the whole ProjectL2 NIC on selected workloads
Tenant actionAttach the ProjectAttach each pod or VM
BGPPlatform-managed FRRNone
Default routeUnchangedUnchanged
Project VPC connected to an approved remote network with BGPWorkloads in Project production use VPC 10.0.0.0/24. The Project VPC router sends approved remote destinations through two managed routing gateway replicas in active and standby mode. Those gateways exchange the Project CIDR and approved prefix 198.51.100.0/24 over eBGP with an external router or firewall. Other traffic continues through the Project's existing default gateway to the Internet. If no managed gateway has a healthy approved route, remote-destination traffic fails closed instead of falling through to the Internet path.eBGPapproved routePROJECT PRODUCTION · VPC 10.0.0.0/24KUBE-DC ROUTINGROUTED DOMAINWorkloadsPods + VMsVPC routerapproved routesdefault separateBGP pairtwo replicasactive · standbyEdge peerrouter / FWmanaged policyRemote CIDR198.51.100.0/24Default pathexisting gatewayInternetSNAT unchangedSeparate paths by destinationApproved remote destinations fail closed; other traffic keeps the default Internet path.
A whole Project VPC—not a second workload interface—reaches only operator-approved remote prefixes through redundant managed gateways. BGP remains platform-managed, while the Project's default Internet path stays unchanged.

Routed Networks are an alpha feature and are disabled by default. The chart installs their additive CRDs, but the manager registers no watches, admission handler, or data-plane reconciliation until the feature gate is enabled and all four CRDs are discoverable.

Security and traffic contract

  • The platform owns the fabric, peer addresses, ASNs, authentication, and import/export policy.
  • An Organization administrator may attach only an allocation delegated to that Organization, and only to one of its Projects.
  • Exports are derived from the authoritative Project VPC CIDR. Tenants cannot supply an advertised prefix or FRR configuration.
  • Imports are explicit. 0.0.0.0/0 and any Project, node, Service, platform, transit, or routing-link prefix are rejected.
  • Imports from different allocations attached to the same Project must not overlap, even when the allocations use different fabrics. This prevents two controllers from competing for the same VPC destination and fail-closed guard.
  • A shared routing domain requires non-overlapping Project CIDRs.
  • v1 is routed-egress: Project-initiated flows and their replies are allowed; externally initiated sessions and Project transit are denied by nftables.
  • The Project's existing Internet/CGNAT default route is never changed.

For every approved destination, the controller owns a priority-30500 Kube-OVN destination guard above FIP and service-LB source reroutes. The guard is allow while destination steering is healthy. Before changing or withdrawing steering, it becomes drop; with zero healthy gateways it stays drop, so a FIP workload cannot send the corporate destination through its Internet gateway. OVN evaluates policy before static routes, which is why the drop is a transition/unhealthy state rather than permanently armed.

Prerequisites

Before creating a fabric, verify:

  1. The existing Kube-OVN ProviderNetwork is Ready on at least two nodes.
  2. The switch trunks the chosen transit VLAN to those nodes.
  3. The external router owns the transit gateway address and is configured for the declared peer ASN.
  4. The transit addresses are reserved exclusively for routing gateways. On a dedicated VLAN, size the CIDR for two addresses per attached Project, plus the external router and network reservations. When peering on an existing platform VLAN, exclude the reserved gateway addresses from that VLAN's existing Subnet IPAM before creating the fabric.
  5. Imported destinations do not overlap any platform or tenant range.
  6. The external routing domain has a non-overlapping prefix for every Project that will attach.
  7. At least two eligible nodes can satisfy required anti-affinity.
  8. Every eligible node permits the namespaced but unsafe net.ipv4.ip_forward pod sysctl. For RKE2, pass kubelet-arg: ["allowed-unsafe-sysctls=net.ipv4.ip_forward"] in the node config and restart RKE2 before enabling the feature. The gateway deliberately fails startup when its pod network namespace cannot forward IPv4.

Use documentation ranges in examples and source control. Never commit a real peer address or TCP-MD5 password to a public repository.

Enable the feature

Set the chart value only on the intended validation or production cluster:

routedNetwork:
enabled: true
backend: frr-project-gateway
namespace: kube-dc-routing
gatewayImage: shalb/kube-dc-routing-gateway:v0.1.13
frrImage: quay.io/frrouting/frr:10.4.1
routingLinkPool: 100.65.0.0/16

The corresponding manager environment gate is ROUTED_NETWORK_ENABLED=true. The chart requires redundant manager replicas when the gate is on and creates the privileged kube-dc-routing namespace, RBAC, webhook, sibling ValidatingAdmissionPolicy, and retained CRDs.

danger

Do not enable only the manager environment variable. Ship the chart first so all four CRDs and the admission configuration exist. A manager watching a missing CRD exits, which can also take unrelated fail-closed webhooks down.

Gate OFF is intentionally inert: no routed controller or webhook is registered, no gateway is created, and no VPC route is programmed.

Create a RoutingFabric

A fabric describes one provider VLAN and its platform-owned BGP policy.

apiVersion: kube-dc.com/v1
kind: RoutingFabric
metadata:
name: corporate-edge
spec:
driver: frr-project-gateway
providerNetwork: ext-cloud
vlanId: 2000
transit:
cidrBlock: 192.0.2.0/29
gateway: 192.0.2.1
excludeIps:
- 192.0.2.1
bgp:
localASN: 65002
peers:
- address: 192.0.2.1
asn: 65001
timers:
keepalive: 3
hold: 9
bfd:
enabled: false
minRx: 300
minTx: 300
multiplier: 3
routingDomain:
mode: shared
nodeSelector:
matchLabels:
network.kube-dc.com/routing: "true"
highAvailability:
replicas: 2

If the peer requires TCP-MD5, create a Secret named by spec.bgp.authenticationSecretRef.name in kube-dc-routing. Its required key is password. Supply it through the installation's secret-management workflow; do not put the value in the CR, a ConfigMap, shell history, or Git. The controller injects it only into a mode-0600 runtime FRR file.

kubectl get routingfabric corporate-edge
kubectl get routingfabric corporate-edge \
-o jsonpath='{.status.phase}{"\t"}{.status.transitAddressesFree}{"\n"}'
kubectl get routingfabric corporate-edge \
-o jsonpath='{range .status.conditions[*]}{.type}{"="}{.status}{" "}{.reason}{"\n"}{end}'

Ready requires the provider network and the selected nodes to be available. For a dedicated tag, the controller creates and owns the provider VLAN, transit Subnet, and NAD. If the same provider and VLAN ID already have exactly one healthy platform-owned Kube-OVN Vlan, the controller references that wire without adopting, mutating, or deleting it. The routed address slice then uses a fabric-owned transit VPC so its IPAM can overlap the platform Subnet without sharing allocations. Multiple matching Vlan objects or a conflicted VLAN fail closed. Do not create aliases for generated objects.

Allocate policy to an Organization

Create the allocation from Infrastructure → Routed Networks, or apply a platform-owned manifest:

apiVersion: kube-dc.com/v1
kind: RoutedNetworkAllocation
metadata:
name: corporate-services
spec:
fabricRef: corporate-edge
organization: acme
importPolicy:
allowedPrefixes:
- 198.51.100.0/24
matchMode: exact-or-more-specific
minPrefixLength: 24
maxPrefixLength: 28
maxPrefixes: 100
allowDefaultRoute: false
exportPolicy:
mode: project-vpc-cidrs
maxPrefixes: 16
sharing:
mode: multiple-projects

An empty allowedPrefixes list means import nothing. It is safe and valid. A default route is not a shortcut for “all corporate networks”; enumerate the destinations that the Organization is allowed to reach.

kubectl get routednetworkallocation corporate-services
kubectl get routednetworkallocation corporate-services \
-o jsonpath='{.status.phase}{"\t"}{.status.attachedProjects}{"\n"}'

Allocation names are the Organization-facing handles. fabricRef and organization are immutable. Delete and recreate an unused allocation to move it to another Organization.

Attachment lifecycle

Organization administrators attach from Organization → Networks → Routed Networks. The request is a ProjectRouteAttachment in the Project backing namespace:

apiVersion: kube-dc.com/v1
kind: ProjectRouteAttachment
metadata:
name: corporate
namespace: acme-production
spec:
allocationRef: corporate-services
direction: routed-egress

The ordered controller pipeline is:

  1. revalidate allocation ownership, sharing, and all prefix collisions;
  2. reserve two transit addresses and one hidden routing-link /29 atomically;
  3. create the Project routing link and controller-owned FRR/nftables config;
  4. create two hardened gateway replicas in kube-dc-routing;
  5. wait until every configured BGP peer is Established on a replica;
  6. install destination drop guards above FIP/SvcLB reroutes;
  7. merge healthy next-hop steering into the existing VPC route slices and replace each transition drop with a healthy destination allow; and
  8. publish attachment and gateway status.

Generated gateways have three interfaces: platform eth0 carries metrics and probes, net1 is the hidden Project-VPC routing link, and net2 is the provider transit VLAN. They run without a service account token, with no privilege escalation, and the bounded NET_ADMIN/NET_BIND_SERVICE/NET_RAW/SYS_ADMIN set required by the pinned FRR zebra binary. SETUID/SETGID are present only so zebra and bgpd can complete FRR's normal privilege drop; both daemons then run as the unprivileged frr user. They use the runtime-default seccomp profile, are not privileged, and have no host network, host PID namespace, or host mounts. The routing capability requirement matches FRR 10.4.1 in MetalLB. Required anti-affinity, topology spread, a PodDisruptionBudget, SecurityGroup, and nftables provide defense in depth.

HA and BFD

The two gateway replicas both establish eBGP. Stateful routed-egress uses deterministic active/standby forwarding: the lowest allocated gateway is the VPC next hop, and per-replica MED 0/100/... is derived from that pod's observed routing-link address rather than its independently allocated transit address. The external router therefore selects the same return path even when the two IP pools assign replicas in different orders. When that session fails, the manager and peer promote the next healthy replica. This keeps established-return nftables enforcement valid without sharing conntrack state between pods.

The pinned three-capability gateway runtime runs zebra and bgpd, so routes accepted by the import route-map are load-bearing in the kernel FIB and peer withdrawal removes them. Readiness proves both the accepted BGP routes and their proto bgp entries in the kernel; zebra VTY loss fails liveness so the pod restarts instead of remaining a silent black hole. It does not run bfdd. Even when BFD is requested, status therefore reports degraded-no-bfd and failover follows the BGP hold timer plus cached exporter readiness. The FRR generator and API retain BFD configuration for a future qualified runtime, but operators must never report BFD as active until both the peer and Kube-OVN next-hop path are proven live.

Inspect operation

kubectl -n acme-production get projectrouteattachment corporate -o yaml
kubectl -n acme-production get projectroutinggateway
kubectl -n kube-dc-routing get deploy,pod,pdb \
-l network.kube-dc.com/project-namespace=acme-production

# Platform-only peer and policy inspection
POD=$(kubectl -n kube-dc-routing get pod \
-l network.kube-dc.com/project-namespace=acme-production \
-o jsonpath='{.items[0].metadata.name}')
kubectl -n kube-dc-routing exec "$POD" -- \
su frr -s /bin/sh -c "vtysh -c 'show bgp ipv4 unicast summary'"
kubectl -n kube-dc-routing exec "$POD" -- \
su frr -s /bin/sh -c "vtysh -c 'show route-map'"

# The same session state as exported to Prometheus. The listener is bound to
# the infra PodIP rather than loopback.
POD_IP=$(kubectl -n kube-dc-routing get pod "$POD" \
-o jsonpath='{.status.podIP}')
kubectl -n kube-dc-routing exec "$POD" -- \
wget -qO- "http://$POD_IP:9100/metrics" | grep kube_dc_bgp_session_up

Run vtysh as the frr user: its VTY sockets belong to the frrvty group, so a default root kubectl exec -- vtysh … does not connect in the hardened container even though bgpd and zebra are healthy.

The controller records Ready, PolicyValid, Redundant, BgpEstablished, RoutesProgrammed, and FailClosedArmed. A healthy two-replica attachment is Ready; one replica is Degraded but available; zero replicas means steering is withdrawn while the drop remains.

The main-only Grafana dashboard has UID kube-dc-routed-networks. Alerts are:

  • BGPPeerDown, BGPAllPeersDown, BGPRouteFlapping, BGPMissingImport
  • BGPFibDesynchronized
  • BGPMaxPrefixExceeded, BGPUnexpectedRoute, RouteLeakDetected
  • RoutingGatewayReplicaDown, RoutingGatewayNoRedundancy
  • RoutingGatewayNftCounterCollectorDown
  • RoutingFabricUnavailable, ProjectRouteProgrammingFailed
  • TransitAddressExhausted

The gateway also exports kube_dc_nft_counter_collector_up. A value of 0 means the capability-isolated nftables snapshot helper is unavailable or its snapshot is stale; BGP metrics remain available, but traffic and drop counters must not be treated as current.

Organization notifications use only organization, project, and routed_network; peer and provider topology remain platform-only.

Audit Events include RoutedNetworkAllocated, ProjectRouteAttached, ProjectRouteDetached, ImportPolicyChanged, BGPPeerChanged, and RoutePolicyViolation. Events carry bounded old/new state; authenticated actor and full old/new request data come from the Kubernetes audit log.

Prove the admission boundary

Use server-side dry-run with the tenant identity you intend to grant. These checks must fail for a Project administrator and for cross-Organization use:

kubectl create -f attachment.yaml --dry-run=server \
--as=project-admin@example.test \
--as-group=acme:project-admin

kubectl create -f attachment.yaml --dry-run=server \
--as=org-admin@example.test \
--as-group=another-org:org-admin

An acme:org-admin request may create/delete its own attachment but receives no update or patch permission. The sibling ValidatingAdmissionPolicy also protects generated gateways, Deployments, PDBs, NADs, Secrets, ConfigMaps, SecurityGroups, Subnets, and the shared VPC route slices. Do not weaken or merge it into the unrelated identity-boundary policy.

Drain and delete

Delete the attachment before reclaiming its allocation or fabric. Teardown is fail closed and intentionally takes several reconciles:

  1. replace the healthy destination allow with a drop before removing steering;
  2. scale gateways to zero and wait for pods to disappear;
  3. delete owned config, security, and routing-link resources;
  4. release the atomic address claim;
  5. remove the residual destination guard last; and
  6. remove the finalizer.

Never force-remove the finalizer during ordinary operation. If emergency recovery requires it, inspect and remove the exact UID-owned VPC routes and generated resources first; otherwise stale forwarding state or an address collision can survive the API object.

Disable the feature only after every ProjectRouteAttachment, RoutedNetworkAllocation, ProjectRoutingGateway, and RoutingFabric has completed deletion.

Troubleshooting

SymptomCheck
Fabric UnavailableProviderNetwork ready nodes, node selector, VLAN trunk, transit Subnet
Attachment PolicyInvalidOrganization ownership, sharing mode, exact collision message
Pods running but not Readyvtysh summary, peer ASN/address, VLAN reachability, MD5, timers
Pod rejected with SysctlForbiddenadd net.ipv4.ip_forward to each eligible kubelet's allowedUnsafeSysctls, then restart kubelet
One ready replicaanti-affinity capacity, failed pod, BGP state; routing should remain available
No ready replicasconfirm steering is absent and FailClosedArmed=True before debugging the peer
Unexpected routesinspect rejected-route metric and peer export; never broaden the import list reflexively
Internet changedtreat as a routing incident: the controller must never own 0.0.0.0/0
Address exhaustionreclaim unused attachments or expand the fabric transit CIDR through a planned replacement

Do not hand-edit generated FRR, nftables, VPC routes, or gateway workloads. Reconciliation restores the declared policy, and manual route edits can break the ordering that makes failure safe.