| Internet-Draft | FALCON | September 2026 |
| Song, et al. | Expires 1 April 2027 | [Page] |
This document describes FALCON (FAst Latency and COngestion Notification), a method that allows a traffic source to learn the queuing delay and congestion status of a path with a notification lag no greater than the one-way propagation delay from the congested node back to the source, which is at most about half of the baseline Round-Trip Time (RTT). FALCON combines in-network telemetry and source routing: a forward packet records the path it traverses, and the receiver returns a high-priority packet that is source-routed along the exact reverse of that path, collecting the state of the forward-direction queues as it passes. Where a hop-by-hop flow control mechanism is used, such as Priority-based Flow Control (PFC) or a backpressure mechanism for lossless WAN transport, the returned packet can also collect the buffer state that determines when each node throttles its upstream neighbor, allowing the source to act before throttling occurs and to distinguish the root of congestion from nodes that are only throttled by it. The fresher telemetry enables more effective traffic steering and congestion control. FALCON is applicable to data center and wide area networks within a single administrative domain and is designed to be realized with IETF In situ OAM (IOAM) and SRv6, with the extensions identified in this document.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 1 April 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License.¶
Many congestion control (CC) and load balancing (LB) schemes rely on timely path congestion status and/or measurements of packet delay as the basis for rate adjustment or path selection. The motivation and requirements for fast network notifications are described in [I-D.ietf-fann-problem-statement], and an overall architecture is described in [I-D.song-fann-framework].¶
End-to-end feedback mechanisms, whether in-band (e.g., ECN [RFC3168] echoed by the receiver, or in-network telemetry echoed in acknowledgments as in [HPCC]) or out-of-band (e.g., mesh probing), deliver the state of a network node to the traffic source only after the carrying packet has completed the remainder of the forward path and the entire return path. The information is therefore at least one baseline RTT old for the first hop and, because RTT grows with queuing, becomes staler as the path becomes more congested, which is exactly when fresh information is most needed. For example, when a packet experiences congestion and is ECN-marked, the congestion itself delays the delivery of the mark to the source. In a Wide Area Network (WAN), where the baseline RTT can be tens to hundreds of milliseconds, a reaction based on such stale feedback may be ineffective or even counterproductive.¶
The root cause of the staleness is that the state is sampled on the forward direction and must then travel the remaining forward path plus the full reverse path. FALCON instead samples the state on the reverse direction: the receiver returns a packet that is source-routed along the exact reverse of the forward path and that reads, at each node, the state of the queue that the forward traffic uses. The lag between sampling at a node and delivery to the source is then only the one-way delay from that node to the source.¶
FALCON is designed to be realized with IOAM trace options [RFC9197] and SRv6 [RFC8754] [RFC8986]. Section 6.2 identifies the extensions that are needed to do so.¶
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.¶
A source S needs to sense the current status of the path toward a destination (e.g., per-hop queuing delay, congestion, total path latency) in order to react in time. For state that originates at a node S_i, the staleness at S cannot be smaller than the one-way delay from S_i to S. For the most distant node on a path with symmetric propagation delays, this lower bound is half the baseline RTT. It is usually well below half of the measured RTT, which additionally includes queuing delay.¶
FALCON aims to approach this lower bound. Its design goals are:¶
Figure 1 illustrates the operation.¶
S S_1 S_2 ... S_n R
| | | | |
(1) |-- P --->|-- P --->|---- ... ----->|-- P --->|
| INT: record node ID and interfaces at S_i |
| | | | |
(2) | | | | R builds P':
| | | | SRH = reverse
| | | | of recorded path
| | | | |
(3) |<-- P' --|<-- P' --|<--- ... ------|<-- P' --|
| high priority; at each S_i read the |
| Forward Queue (egress queue of the |
| interface on which P' arrived) and, with |
| HFC, the Flow-Control Buffer of the |
| interface through which P' leaves |
| | | | |
(4) S computes per-hop/path queuing delay and congestion
Because each S_i is sampled when P' passes it, the staleness of the state of S_i at S is the one-way delay from S_i to S. The state of the congested node that matters most to S is therefore delivered as quickly as any notification originating at that node could be, without requiring transit nodes to generate packets.¶
P carries an INT trace instruction requesting, at each FALCON transit node, a node identifier and the ingress and egress interface identifiers. With IOAM [RFC9197], this corresponds to the node_id and the ingress_if_id/egress_if_id data fields of the IOAM trace option. P also indicates the traffic class of the monitored traffic if it differs from the traffic class in which P itself is forwarded.¶
If P is a data packet, S SHOULD apply the recording instruction only to a sample of packets (Section 7.1). If P is a dedicated probe, it MUST be forwarded along the same path as the monitored flow, e.g., by using the same ECMP-relevant header fields or the same SR segment list.¶
If S already steers the monitored traffic with an explicit SR segment list whose segments identify every hop, recording the path is unnecessary; R can derive the return path from that segment list, or S can itself provide the return segment list in P.¶
R maps each recorded (node identifier, interface identifier) pair to an SRv6 SID. To guarantee that P' arrives at S_i on the interface through which P left S_i, the segment list SHOULD use, for each hop, an adjacency SID (e.g., End.X [RFC8986]) of S_(i+1) for the link between S_(i+1) and S_i, where S_(i+1) is identified by the recorded node identifier and the link by the recorded ingress interface of S_(i+1). Node SIDs alone are insufficient where parallel links or ECMP exist between adjacent nodes. The final segment identifies S.¶
The mapping from IOAM node and interface identifiers to SIDs is provisioned by the operator or learned via the control plane; alternatively, where the deployment permits it, the transit nodes MAY record the relevant SID directly. The mechanism for distributing this mapping is outside the scope of this version of the document.¶
The length of the segment list grows with the hop count. To bound the overhead, particularly in the WAN, compressed SIDs [RFC9800] SHOULD be used where supported.¶
P' MUST be marked with a DSCP [RFC2474] that maps to a high-priority, low-latency per-hop behavior within the FALCON domain, and it carries an INT instruction requesting the Forward Queue data of Section 5.3, together with the monitored traffic class. P' MAY be piggybacked on a transport-layer acknowledgment or congestion notification packet that R would send to S anyway.¶
When P' arrives at S_i on interface r, S_i reports the state of the Forward Queue, i.e., the egress queue of interface r for the monitored traffic class. Depending on the capability of S_i and the requested data, S_i reports one or more of:¶
Note that the existing IOAM queue depth data field [RFC9197] reports the egress queue of the interface through which the IOAM packet itself leaves the node, which is not the Forward Queue. A new data field is therefore needed (Section 6.2).¶
The congestion indication is carried as telemetry data in P'. It MUST NOT be signaled by setting the ECN field of the IP header of P', since that field reflects congestion experienced by P' itself.¶
In lossless networks, a node S_i protects its buffer by sending a Throttle Signal to its upstream neighbor S_(i-1) when the occupancy of its Flow-Control Buffer exceeds the Trigger Threshold, and by releasing it when the occupancy falls sufficiently. Examples are PFC in RoCEv2 data center fabrics, credit-based link-level flow control, and hop-by-hop backpressure applied per slice or per flow over a WAN [I-D.han-rtgwg-fine-grained-backpressure]. The mechanisms differ in the signal they use and in the granularity of the flow-control class, but share the same structure: a per-class buffer accounting on the receiving side of a link, a threshold, and a signal that throttles the sending side. The Flow-Control Buffer is accounted separately from the Forward Queue, and its Trigger Threshold may be dynamic. The depth of the Forward Queue therefore does not by itself indicate how close S_i is to throttling its upstream neighbor.¶
When P' leaves S_i through interface q, S_i MAY report, for the flow-control class of the monitored traffic, one or more of:¶
The reported information is expressed independently of the specific HFC mechanism, so that the source can interpret it uniformly across a path that may combine, for example, PFC within data centers and a different backpressure mechanism on the interconnecting WAN. These data are sampled when P' passes S_i and are therefore subject to the same staleness bound as the Forward Queue state (Section 5.7). With them, S can:¶
HFC reacts hop by hop, within the headroom of a single link. FALCON does not replace it: a burst shorter than the control loop between S and S_i can still trigger throttling. The purpose of this data is to let S reduce sustained buildup that would otherwise lead to repeated or prolonged throttling. This benefit is greatest where the headroom per link is costly relative to the link delay, as on long-distance WAN links.¶
In some devices, flow-control buffer accounting and throttle state are kept in the buffer manager and are not directly accessible while a packet of a different interface is processed. A node that cannot provide a requested value SHOULD mark it as unavailable rather than omit it.¶
If transit nodes can perform simple arithmetic in the data plane, P' MAY carry aggregate values instead of a per-hop list: the running sum of Forward Queue delay (yielding the path queuing delay), or the maximum Forward Queue depth or delay together with the identifier of the node reporting it (yielding the bottleneck), updated by a compare-and-replace operation. Where HFC state is collected, P' MAY similarly carry the minimum headroom to the Trigger Threshold along the path with the identifier of the node reporting it, and a flag set by any node that is throttling or being throttled. Aggregation keeps the size of P' constant regardless of path length.¶
If transit nodes cannot aggregate, P' collects per-hop values, and S performs the computation using its knowledge of link capacities.¶
The delay of P over the path consists of a baseline component (propagation, serialization, and minimum forwarding delay) and a queuing component. The queuing component is the sum of the Forward Queue delays reported by P'. Where a node reports only queue depth, S estimates the queuing delay at S_i as the queue depth divided by the service rate of that queue. The service rate equals the link rate only when the monitored class is the sole active class on the link and is not subject to flow control (e.g., HFC throttling, see Section 5.4) or shaping; S SHOULD prefer node-reported delay estimates when available.¶
The baseline component can be obtained by the source as the minimum observed RTT (or minimum P-to-P' round trip) over a suitable window, divided between directions under an assumption of symmetric propagation delay; or, where clocks are synchronized, from IOAM timestamps carried in P and P'. The estimated current path delay of P is the baseline component plus the collected queuing component.¶
S passes the per-hop state, path queuing delay, bottleneck information, and congestion indications to its CC or LB function. How S reacts is outside the scope of this document.¶
Let d(X,Y) denote the one-way delay from X to Y without queuing, and Q_fwd(i) and Q_rev the queuing delays encountered after S_i on the forward path and on the reverse path from R to S, respectively. For the state of S_i:¶
The gain is largest for nodes close to the source and grows with the amount of forward-path queuing. Since a reaction by S takes effect at S_i only after d(S,S_i), the FALCON control loop for a bottleneck at S_i is approximately the baseline RTT between S and S_i, which is the physical lower bound for any source-based reaction.¶
Staleness is distinct from sampling frequency: the interval between consecutive P' packets is determined by the sampling rate of P (Section 7.1). In addition, the per-hop samples carried by one P' are taken at slightly different times (S_n first, S_1 last), spread over d(S_n,S_1).¶
In IPv6 networks, P can carry the IOAM trace option in an IPv6 Hop-by-Hop Options header [RFC9486]. For P', which carries both an SRH and an IOAM trace, two approaches are possible:¶
The following extensions are needed for a standards-based realization:¶
Generating a P' for every packet is unnecessary and costly at high link rates. S SHOULD request recording on a configurable subset of packets, e.g., once per configured time interval, per configured number of bytes, or per RTT per flow. R MUST rate-limit the generation of P', and the aggregate volume of P' in the high-priority class MUST be bounded so that it cannot starve other traffic in that class.¶
The trace data added to P can cause P to exceed the path MTU. S SHOULD account for the maximum trace size when selecting data packets to carry the instruction, or use dedicated probes.¶
The HFC data of Section 5.4 are useful only where an HFC mechanism is enabled for the monitored traffic; S SHOULD NOT request them otherwise.¶
Nodes that do not support FALCON do not record themselves in P and are therefore not included in the return segment list; the link state between two FALCON nodes separated by a non-capable node cannot be observed, and P' traverses the intervening segment using normal forwarding, which may not follow the reverse of the forward path. S SHOULD detect gaps (e.g., via hop-count or TTL information in the trace) and treat the corresponding portion of the path as unobserved.¶
P' reflects the path taken by the specific P from which it was generated. If the monitored flow is re-hashed or re-routed, S SHOULD discard state for the old path. Where a link is a member of a link aggregation group, the Forward Queue is per member link; the recorded interface identifier SHOULD identify the member link. FALCON requires only that each link be bidirectional; it does not require symmetric routing, because P' is explicitly steered. Estimates of the baseline delay that assume symmetric propagation are affected by asymmetric link delays.¶
FALCON is intended for use within a single administrative domain that is both an SR domain and an IOAM domain, consistent with the deployment scope of [I-D.song-fann-framework]. In a data center network the paths are short and the main benefit is sub-RTT reaction to incast and transient congestion. In a WAN the baseline RTT is large, and the absolute reduction in staleness is correspondingly large.¶
The security considerations of IOAM [RFC9197] [RFC9486] and SRv6 [RFC8754] [RFC8986] apply. In particular, FALCON packets MUST be confined to a domain that is both an SR domain and an IOAM domain; packets carrying FALCON instructions, SRHs, or the P' traffic class MUST be filtered at the domain boundary.¶
Additional threats specific to FALCON include:¶
This document has no IANA actions at this time. Future versions are expected to request allocations for the extensions listed in Section 6.2, e.g., new bits in the IOAM Trace-Type registry and any flags or code points required for the P' request indication.¶