Troubleshooting AWS Direct Connect Link Down and BGP States
AWS Direct Connect provides a private connection between your network and AWS, but troubleshooting it requires separating physical connectivity from VLAN configuration and BGP routing. A link can appear healthy at the carrier level while the virtual interface remains down, or the circuit can be operational while the BGP session never reaches Established.
The fastest approach is to follow the dependency chain: physical port, cross-connect, virtual interface, 802.1Q VLAN, IP addressing, BGP authentication, and route policy. Checking these layers in order prevents time being wasted on route advertisements when the underlying VLAN is not passing traffic.
AWS exposes several states in the Direct Connect console, while routers and firewalls report their own interface and BGP conditions. Comparing both sides of the connection is essential, especially when a hosted connection, third-party colocation provider, or redundant Direct Connect location is involved.
Understand The Direct Connect Failure Domains
A dedicated connection links your network to an AWS Direct Connect location through a physical port and a cross-connect. A hosted connection or hosted virtual interface may add a network service provider between your equipment and AWS. Each handoff introduces another point where the circuit can remain administratively configured but operationally unavailable.
Above the physical connection, you configure a private, public, or transit virtual interface. The virtual interface carries an assigned VLAN and uses BGP to exchange routes. A failure at any lower layer affects the higher layers, so a BGP alert may simply be a symptom of a disabled VLAN or unavailable cross-connect.
The Direct Connect console should be treated as one source of evidence rather than the only authority. Check the AWS connection, virtual interface, and BGP status alongside switch port counters, router logs, and carrier monitoring.
Confirm Physical And Circuit State
Begin with the Direct Connect connection state. If AWS reports the connection as down, inspect the local router or switch interface for link, speed, duplex, optics, and error conditions. A dark fiber, transceiver mismatch, incorrect port assignment, or provider maintenance event can prevent the connection from becoming available.
On the local device, look for loss of signal, input errors, CRC errors, frequent link flaps, or an unexpected negotiated speed. Confirm that the optic type and wavelength match the service specification. If the port is up but errors continue increasing, collect interface diagnostics before repeatedly resetting the connection.
For a hosted service, ask the provider to verify the service handoff, VLAN presentation, and cross-connect status. Record timestamps for every state change. AWS Support or the connectivity provider can correlate those times with port alarms and maintenance records.
Validate The VLAN And Virtual Interface
A private virtual interface depends on the correct VLAN tag reaching the AWS-facing port. Verify that the router uses the VLAN ID assigned to the virtual interface and that intermediate switches permit the tag. Native VLAN behavior, incorrect trunk configuration, or a mismatched subinterface can leave the physical link up while the VIF remains down.
Check the IP addresses configured on both sides of the BGP peering. The customer router must use the customer BGP address, while the AWS peer address must match the values shown in the Direct Connect console. Confirm the subnet mask and ensure no duplicate address exists elsewhere in the environment.
The virtual interface type also matters. A private VIF commonly connects to a virtual private gateway or Direct Connect gateway, while a transit VIF connects to a Direct Connect gateway associated with a transit gateway. A correct BGP session does not guarantee reachability if the VIF is attached to the wrong AWS routing construct.
| Observed State | Likely Area | Useful Checks |
|---|---|---|
| Connection down | Port, optic, cross-connect, or provider circuit | Link LEDs, interface counters, carrier alarms, AWS connection status |
| Connection up, VIF down | VLAN or virtual interface configuration | VLAN ID, trunk allowance, VIF association, subinterface state |
| BGP Idle | Local configuration or disabled neighbor | Peer address, ASN, update source, administrative shutdown |
| BGP Active or Connect | TCP 179 path, ACL, firewall, or addressing | Ping where supported, route lookup, control-plane filters |
| BGP OpenSent | ASN, BGP authentication, or capability mismatch | Local/remote ASN, MD5 key, BGP logs |
| BGP Established, no routes | Policy or AWS attachment | Import/export policy, prefixes, route tables, allowed advertisements |
Interpret BGP Neighbor States
Idle generally indicates that the neighbor is not attempting a session, has been administratively disabled, or is waiting for the local process to retry. Check whether the neighbor is configured under the correct VRF and whether a route exists to the AWS peer address.
Connect means the router is trying to establish the TCP session. Active often indicates that the TCP connection failed and the device is retrying. These states frequently point to a wrong peer IP, missing connected route, ACL blocking TCP port 179, or a control-plane policy that rejects the connection.
OpenSent means TCP succeeded and the router sent a BGP OPEN message, but the peers have not completed negotiation. Check the local and remote autonomous system numbers, BGP authentication, router ID behavior, and supported capabilities. A password mismatch can prevent the session from progressing even though the interface and IP addresses are correct.
OpenConfirm indicates that the OPEN exchange succeeded and the routers are waiting for the final keepalive. If the neighbor repeatedly cycles through this state, inspect BGP timers, packet loss, and stateful firewall behavior. Established confirms that the control-plane session is active, but route exchange still requires correct policy.
Check ASN, Authentication And Routing Policy
AWS Direct Connect requires the customer-side BGP ASN and the AWS-side ASN to match the values configured for the virtual interface. Review whether the router uses a public or private ASN as expected, and check for inherited neighbor-group settings that override the intended values.
If BGP MD5 authentication is enabled, compare the key exactly, including capitalization and special characters. Some platforms hide the configured password after entry, so reapplying the known value on both sides may be quicker than trying to inspect it.
Once the session is Established, inspect advertised and received routes. Confirm that the on-premises prefixes are permitted for export and that AWS routes are accepted by the inbound policy. Prefix limits, route filters, maximum-prefix settings, and accidental community-based policies can produce a healthy BGP session with no usable connectivity.
Use A Repeatable Recovery Sequence
A disciplined workflow reduces unnecessary resets and creates evidence for escalation. Capture the current state before changing configuration, including console screenshots, router outputs, interface counters, BGP logs, and timestamps.
Work from the lowest dependency to the highest. Do not troubleshoot route maps while the virtual interface is down, and do not replace optics when the physical port is healthy and the BGP log clearly reports an ASN mismatch.
Recommended checks include:
- Verify the AWS connection and virtual interface status.
- Confirm local port, optic, cross-connect, and carrier health.
- Validate VLAN tagging, subinterface state, and VIF attachment.
- Compare BGP peer IPs, ASNs, authentication keys, and VRF context.
- Review TCP 179 filtering, route policy, prefix limits, and advertised routes.
After each change, allow the relevant timers to expire and record whether the state advances. If the link remains down after local checks, provide AWS or the carrier with the connection ID, virtual interface ID, device logs, interface counters, and precise failure timestamps.
Build the same validation into monitoring for production circuits. Alert separately on physical link state, VIF state, BGP Established status, and route count so that a routing outage is not confused with a fiber or provider failure.
Use this layered process during the next AWS Direct Connect incident, preserve the evidence from each checkpoint, and document the working configuration for every redundant circuit and BGP peer.