<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE rfc [
  <!ENTITY nbsp    "&#160;">
  <!ENTITY zwsp   "&#8203;">
  <!ENTITY nbhy   "&#8209;">
  <!ENTITY wj     "&#8288;">
]>

<?xml-stylesheet type="text/xsl" href="rfc2629.xslt"?>

<rfc
  xmlns:xi="http://www.w3.org/2001/XInclude"
  category="info"
  docName="draft-song-fann-falcon-00"
  ipr="trust200902"
  submissionType="IETF"
  consensus="true"
  xml:lang="en"
  tocInclude="true"
  tocDepth="3"
  symRefs="true"
  sortRefs="true"
  version="3">

  <!-- ============================================================ -->
  <!--  FRONT MATTER                                                -->
  <!-- ============================================================ -->
  <front>
    <title abbrev="FALCON">
      FALCON: Fast Latency and Congestion Notification Using Reverse-Path In-Network Telemetry
    </title>

    <seriesInfo name="Internet-Draft" value="draft-song-fann-falcon-00"/>

    <author fullname="Haoyu Song" initials="H." surname="Song" role="editor">
      <organization>Futurewei Technologies</organization>
      <address>
        <postal>
          <country>US</country>
        </postal>
        <email>haoyu.song@futurewei.com</email>
      </address>
    </author>

    <author fullname="Ying Wan" initials="Y." surname="Wan">
      <organization>Southeast University</organization>
      <address>
        <postal>
          <country>CN</country>
        </postal>
        <email>wy25@seu.edu.cn</email>
      </address>
    </author>

    <author fullname="Keyi Zhu" initials="K." surname="Zhu">
      <organization>Huawei Technologies</organization>
      <address>
        <postal>
          <country>CN</country>
        </postal>
        <email>zhukeyi@huawei.com</email>
      </address>
    </author>

    <date/>

    <area>Routing</area>
    <workgroup>Fast Network Notifications</workgroup>

    <keyword>FANN</keyword>
    <keyword>IOAM</keyword>
    <keyword>SRv6</keyword>
    <keyword>congestion control</keyword>
    <keyword>load balancing</keyword>

    <abstract>
      <t>
        This document describes FALCON (FAst Latency and COngestion
        Notification), a method that allows a traffic source to learn the
        queuing delay and congestion status of a path with a notification
        lag no greater than the one-way propagation delay from the
        congested node back to the source, which is at most about half of
        the baseline Round-Trip Time (RTT). FALCON combines in-network
        telemetry and source routing: a forward packet records the path it
        traverses, and the receiver returns a high-priority packet that is
        source-routed along the exact reverse of that path, collecting the
        state of the forward-direction queues as it passes. Where a
        hop-by-hop flow control mechanism is used, such as Priority-based
        Flow Control (PFC) or a backpressure mechanism for lossless WAN
        transport, the returned packet can also collect the buffer state
        that determines when each node throttles its upstream neighbor,
        allowing the source to act before throttling occurs and to
        distinguish the root of congestion from nodes that are only
        throttled by it. The fresher
        telemetry enables more effective traffic steering and congestion
        control. FALCON is applicable to data center and wide area
        networks within a single administrative domain and is designed to
        be realized with IETF In situ OAM (IOAM) and SRv6, with the
        extensions identified in this document.
      </t>
    </abstract>

  </front>

  <!-- ============================================================ -->
  <!--  MIDDLE                                                      -->
  <!-- ============================================================ -->
  <middle>

    <!-- 1. Introduction -->
    <section anchor="intro">
      <name>Introduction</name>
      <t>
        Many congestion control (CC) and load balancing (LB) schemes rely
        on timely path congestion status and/or measurements of packet
        delay as the basis for rate adjustment or path selection. The
        motivation and requirements for fast network notifications are
        described in <xref target="I-D.ietf-fann-problem-statement"/>, and
        an overall architecture is described in
        <xref target="I-D.song-fann-framework"/>.
      </t>
      <t>
        End-to-end feedback mechanisms, whether in-band (e.g., ECN
        <xref target="RFC3168"/> echoed by the receiver, or in-network
        telemetry echoed in acknowledgments as in <xref target="HPCC"/>)
        or out-of-band (e.g., mesh probing), deliver the state of a
        network node to the traffic source only after the carrying packet
        has completed the remainder of the forward path and the entire
        return path. The information is therefore at least one baseline
        RTT old for the first hop and, because RTT grows with queuing,
        becomes staler as the path becomes more congested, which is
        exactly when fresh information is most needed. For example, when
        a packet experiences congestion and is ECN-marked, the congestion
        itself delays the delivery of the mark to the source. In a Wide
        Area Network (WAN), where the baseline RTT can be tens to hundreds
        of milliseconds, a reaction based on such stale feedback may be
        ineffective or even counterproductive.
      </t>
      <t>
        The root cause of the staleness is that the state is sampled on
        the forward direction and must then travel the remaining forward
        path plus the full reverse path. FALCON instead samples the state
        on the reverse direction: the receiver returns a packet that is
        source-routed along the exact reverse of the forward path and that
        reads, at each node, the state of the queue that the forward
        traffic uses. The lag between sampling at a node and delivery to
        the source is then only the one-way delay from that node to the
        source.
      </t>
      <t>
        FALCON is designed to be realized with IOAM trace options
        <xref target="RFC9197"/> and SRv6 <xref target="RFC8754"/>
        <xref target="RFC8986"/>. <xref target="gaps"/> identifies the
        extensions that are needed to do so.
      </t>
    </section>

    <!-- 2. Terminology -->
    <section anchor="terminology">
      <name>Terminology</name>

      <section anchor="requirements-language">
        <name>Requirements Language</name>
        <t>
          The key words "<bcp14>MUST</bcp14>", "<bcp14>MUST NOT</bcp14>",
          "<bcp14>REQUIRED</bcp14>", "<bcp14>SHALL</bcp14>",
          "<bcp14>SHALL NOT</bcp14>", "<bcp14>SHOULD</bcp14>",
          "<bcp14>SHOULD NOT</bcp14>", "<bcp14>RECOMMENDED</bcp14>",
          "<bcp14>NOT RECOMMENDED</bcp14>", "<bcp14>MAY</bcp14>", and
          "<bcp14>OPTIONAL</bcp14>" in this document are to be interpreted
          as described in BCP 14 <xref target="RFC2119"/>
          <xref target="RFC8174"/> when, and only when, they appear in
          all capitals, as shown here.
        </t>
      </section>

      <section anchor="definitions">
        <name>Definitions</name>
        <dl newline="true" spacing="normal">
          <dt>INT:</dt>
          <dd>In-Network Telemetry. A data packet or a dedicated probe
            packet carries an instruction header that causes network nodes
            on its path to add node data to the packet. The IOAM trace
            options <xref target="RFC9197"/> are an instance of INT.</dd>

          <dt>SR:</dt>
          <dd>Source Routing. The originator of a packet specifies the
            path of the packet as an ordered list of segments. SRv6
            <xref target="RFC8754"/> realizes SR over IPv6.</dd>

          <dt>Source (S):</dt>
          <dd>The node that sends the monitored traffic and consumes FALCON
            notifications. S may be a host, a NIC or DPU, or an ingress
            router/switch that steers traffic.</dd>

          <dt>Receiver (R):</dt>
          <dd>The node that receives the forward packet and generates the
            FALCON Return Packet.</dd>

          <dt>Transit Node (S_i):</dt>
          <dd>The i-th FALCON-capable router or switch on the forward path
            from S to R, i = 1..n.</dd>

          <dt>Forward Packet (P):</dt>
          <dd>A packet from S to R, either a data packet of a monitored
            flow or a dedicated probe, that carries an INT instruction to
            record the path.</dd>

          <dt>FALCON Return Packet (P'):</dt>
          <dd>A packet generated by R in response to P, source-routed along
            the reverse of the path recorded by P, that collects the state
            of the forward-direction queues and delivers it to S.</dd>

          <dt>Forward Queue of S_i:</dt>
          <dd>The egress queue of S_i, for the traffic class of the
            monitored traffic, on the interface through which P left S_i.
            This is the same physical interface on which P' arrives at
            S_i.</dd>

          <dt>Hop-by-Hop Flow Control (HFC):</dt>
          <dd>Any mechanism by which a node, based on its own buffer
            state, signals its upstream neighbor to stop, slow down, or
            resume transmission of a class of traffic. The class
            (the "flow-control class") may be a priority, a slice, a
            tunnel, or a flow, depending on the mechanism. Examples
            include PFC <xref target="IEEE8021Qbb"/>, credit-based
            link-level flow control, and hop-by-hop backpressure for
            lossless transport over WAN
            <xref target="I-D.han-rtgwg-fine-grained-backpressure"/>.</dd>

          <dt>Throttle Signal:</dt>
          <dd>The HFC signal that stops or slows the upstream neighbor,
            e.g., a PFC PAUSE frame, a withheld or reduced credit grant, or
            a backpressure message.</dd>

          <dt>Flow-Control Buffer of S_i:</dt>
          <dd>The buffer accounting of S_i whose state determines when S_i
            sends a Throttle Signal to its upstream neighbor for the
            flow-control class of the monitored traffic, on the interface
            through which P entered S_i (e.g., the ingress priority group
            counter in a shared-buffer switch using PFC, or the per-slice
            buffer used by a WAN backpressure mechanism). This is the same
            physical interface through which P' leaves S_i.</dd>

          <dt>Trigger Threshold:</dt>
          <dd>The occupancy of the Flow-Control Buffer at which S_i sends a
            Throttle Signal, e.g., the XOFF threshold of PFC. It may be
            static or dynamic.</dd>

          <dt>Baseline RTT:</dt>
          <dd>The RTT between S and R excluding queuing delay, i.e., the
            sum of propagation, serialization, and minimum forwarding
            delays in both directions.</dd>

          <dt>Staleness:</dt>
          <dd>For a given piece of node state, the time between the instant
            it is sampled at the node and the instant it is available at
            S.</dd>
        </dl>
      </section>
    </section>

    <!-- 3. Problem Statement -->
    <section anchor="problem">
      <name>Problem Statement and Design Goals</name>
      <t>
        A source S needs to sense the current status of the path toward a
        destination (e.g., per-hop queuing delay, congestion, total path
        latency) in order to react in time. For state that originates at a
        node S_i, the staleness at S cannot be smaller than the one-way
        delay from S_i to S. For the most distant node on a path with
        symmetric propagation delays, this lower bound is half the
        baseline RTT. It is usually well below half of the measured RTT,
        which additionally includes queuing delay.
      </t>
      <t>
        FALCON aims to approach this lower bound. Its design goals are:
      </t>
      <dl newline="false" spacing="normal">
        <dt>Accurate:</dt>
        <dd>The notification reflects the state of the queues actually
          used by the monitored traffic on the path in question.</dd>
        <dt>Timely:</dt>
        <dd>The staleness of the state of each node approaches the one-way
          delay from that node to the source.</dd>
        <dt>Lightweight:</dt>
        <dd>The mechanism imposes low packet and bandwidth overhead,
          requires no packet generation at transit nodes, and needs no
          per-flow state in the network.</dd>
        <dt>Reuse of existing standards:</dt>
        <dd>The mechanism is built on existing IETF data-plane protocols,
          with the minimum extensions needed.</dd>
      </dl>
    </section>

    <!-- 4. Solution Overview -->
    <section anchor="overview">
      <name>Solution Overview</name>
      <t>
        <xref target="fig-overview"/> illustrates the operation.
      </t>
      <figure anchor="fig-overview">
        <name>FALCON Operation</name>
        <artwork type="ascii-art"><![CDATA[
     S        S_1       S_2     ...     S_n        R
     |         |         |               |         |
 (1) |-- P --->|-- P --->|---- ... ----->|-- P --->|
     |  INT: record node ID and interfaces at S_i  |
     |         |         |               |         |
 (2) |         |         |               |  R builds P':
     |         |         |               |  SRH = reverse
     |         |         |               |  of recorded path
     |         |         |               |         |
 (3) |<-- P' --|<-- P' --|<--- ... ------|<-- P' --|
     |  high priority; at each S_i read the        |
     |  Forward Queue (egress queue of the         |
     |  interface on which P' arrived) and, with   |
     |  HFC, the Flow-Control Buffer of the        |
     |  interface through which P' leaves          |
     |         |         |               |         |
 (4) S computes per-hop/path queuing delay and congestion
]]></artwork>
      </figure>
      <ol spacing="normal">
        <li>S sends a Forward Packet P (a data packet of flow F or a
          dedicated probe) carrying an INT instruction that causes each
          transit node S_i to record its node identifier and the
          identifiers of the interfaces through which P entered and left
          it.</li>
        <li>On receiving P, R constructs a FALCON Return Packet P'. P'
          carries a segment list that steers it through S_n, ..., S_2,
          S_1 to S, pinned to the same links P used (<xref target="return-path"/>),
          and an INT instruction to collect Forward Queue state.</li>
        <li>P' is sent in a high-priority traffic class, so it experiences
          negligible queuing on the return path. When P' arrives at S_i
          on interface r, P left S_i through the same interface r, so S_i
          reports the state of the egress queue of r for the traffic class
          of the monitored traffic. Where HFC is deployed, S_i also
          reports the state of the Flow-Control Buffer, i.e., of the
          interface q through which P' leaves S_i and through which P
          entered it (<xref target="hfc"/>).</li>
        <li>S uses the collected data to derive per-hop and path queuing
          delay, the location and severity of congestion, and the total
          path delay (<xref target="source-proc"/>), and feeds these to its
          CC or LB function.</li>
      </ol>
      <t>
        Because each S_i is sampled when P' passes it, the staleness of
        the state of S_i at S is the one-way delay from S_i to S. The
        state of the congested node that matters most to S is therefore
        delivered as quickly as any notification originating at that node
        could be, without requiring transit nodes to generate packets.
      </t>
    </section>

    <!-- 5. Detailed Operation -->
    <section anchor="spec">
      <name>Detailed Operation</name>

      <section anchor="fwd-path">
        <name>Forward Path Recording</name>
        <t>
          P carries an INT trace instruction requesting, at each FALCON
          transit node, a node identifier and the ingress and egress
          interface identifiers. With IOAM <xref target="RFC9197"/>, this
          corresponds to the node_id and the ingress_if_id/egress_if_id
          data fields of the IOAM trace option. P also indicates the
          traffic class of the monitored traffic if it differs from the
          traffic class in which P itself is forwarded.
        </t>
        <t>
          If P is a data packet, S <bcp14>SHOULD</bcp14> apply the
          recording instruction only to a sample of packets
          (<xref target="sampling"/>). If P is a dedicated probe, it
          <bcp14>MUST</bcp14> be forwarded along the same path as the
          monitored flow, e.g., by using the same ECMP-relevant header
          fields or the same SR segment list.
        </t>
        <t>
          If S already steers the monitored traffic with an explicit SR
          segment list whose segments identify every hop, recording the
          path is unnecessary; R can derive the return path from that
          segment list, or S can itself provide the return segment list in
          P.
        </t>
      </section>

      <section anchor="return-path">
        <name>Return Packet Construction</name>
        <t>
          R maps each recorded (node identifier, interface identifier)
          pair to an SRv6 SID. To guarantee that P' arrives at S_i on the
          interface through which P left S_i, the segment list
          <bcp14>SHOULD</bcp14> use, for each hop, an adjacency SID (e.g.,
          End.X <xref target="RFC8986"/>) of S_(i+1) for the link between
          S_(i+1) and S_i, where S_(i+1) is identified by the recorded
          node identifier and the link by the recorded ingress interface
          of S_(i+1). Node SIDs alone are insufficient where parallel
          links or ECMP exist between adjacent nodes. The final segment
          identifies S.
        </t>
        <t>
          The mapping from IOAM node and interface identifiers to SIDs is
          provisioned by the operator or learned via the control plane;
          alternatively, where the deployment permits it, the transit
          nodes MAY record the relevant SID directly. The mechanism for
          distributing this mapping is outside the scope of this
          version of the document.
        </t>
        <t>
          The length of the segment list grows with the hop count. To bound
          the overhead, particularly in the WAN, compressed SIDs
          <xref target="RFC9800"/> <bcp14>SHOULD</bcp14> be used where
          supported.
        </t>
        <t>
          P' <bcp14>MUST</bcp14> be marked with a DSCP
          <xref target="RFC2474"/> that maps to a high-priority,
          low-latency per-hop behavior within the FALCON domain, and it
          carries an INT instruction requesting the Forward Queue data of
          <xref target="per-hop"/>, together with the monitored traffic
          class. P' MAY be piggybacked on a transport-layer acknowledgment
          or congestion notification packet that R would send to S anyway.
        </t>
      </section>

      <section anchor="per-hop">
        <name>Per-Hop Processing on the Return Path</name>
        <t>
          When P' arrives at S_i on interface r, S_i reports the state of
          the Forward Queue, i.e., the egress queue of interface r for the
          monitored traffic class. Depending on the capability of S_i and
          the requested data, S_i reports one or more of:
        </t>
        <ul spacing="normal">
          <li>the Forward Queue depth, in bytes;</li>
          <li>an estimate of the Forward Queue queuing delay, preferably
            derived from hardware sojourn-time measurement or the recent
            dequeue rate of the queue;</li>
          <li>a congestion indication, set when the Forward Queue depth or
            delay exceeds a configured threshold (e.g., the ECN marking
            threshold of that queue);</li>
          <li>the transmit utilization of interface r.</li>
        </ul>
        <t>
          Note that the existing IOAM queue depth data field
          <xref target="RFC9197"/> reports the egress queue of the
          interface through which the IOAM packet itself leaves the node,
          which is not the Forward Queue. A new data field is therefore
          needed (<xref target="gaps"/>).
        </t>
        <t>
          The congestion indication is carried as telemetry data in P'. It
          <bcp14>MUST NOT</bcp14> be signaled by setting the ECN field of
          the IP header of P', since that field reflects congestion
          experienced by P' itself.
        </t>
      </section>

      <section anchor="hfc">
        <name>Hop-by-Hop Flow Control State</name>
        <t>
          In lossless networks, a node S_i protects its buffer by sending
          a Throttle Signal to its upstream neighbor S_(i-1) when the
          occupancy of its Flow-Control Buffer exceeds the Trigger
          Threshold, and by releasing it when the occupancy falls
          sufficiently. Examples are PFC in RoCEv2 data center fabrics,
          credit-based link-level flow control, and hop-by-hop
          backpressure applied per slice or per flow over a WAN
          <xref target="I-D.han-rtgwg-fine-grained-backpressure"/>. The
          mechanisms differ in the signal they use and in the granularity
          of the flow-control class, but share the same structure: a
          per-class buffer accounting on the receiving side of a link,
          a threshold, and a signal that throttles the sending side. The
          Flow-Control Buffer is accounted separately from the Forward
          Queue, and its Trigger Threshold may be dynamic. The depth of
          the Forward Queue therefore does not by itself indicate how
          close S_i is to throttling its upstream neighbor.
        </t>
        <t>
          When P' leaves S_i through interface q, S_i
          <bcp14>MAY</bcp14> report, for the flow-control class of the
          monitored traffic, one or more of:
        </t>
        <ul spacing="normal">
          <li>the Flow-Control Buffer occupancy, in bytes;</li>
          <li>the headroom remaining before the Trigger Threshold
            currently in effect, or the occupancy as a fraction of that
            threshold, which is comparable across nodes with different
            buffer sizes and threshold settings;</li>
          <li>an indication that S_i is currently throttling its upstream
            neighbor on interface q;</li>
          <li>an indication that the Forward Queue is currently throttled
            by S_(i+1) on interface r, optionally with the throttled
            duration or the number of Throttle Signals received in a
            recent interval.</li>
        </ul>
        <t>
          The reported information is expressed independently of the
          specific HFC mechanism, so that the source can interpret it
          uniformly across a path that may combine, for example, PFC
          within data centers and a different backpressure mechanism on
          the interconnecting WAN. These data are sampled when P' passes
          S_i and are therefore subject to the same staleness bound as
          the Forward Queue state (<xref target="freshness"/>). With them,
          S can:
        </t>
        <ul spacing="normal">
          <li>Act before throttling occurs. A shrinking headroom is an
            early warning that lets S reduce its rate or move traffic
            before a Throttle Signal is sent, reducing the frequency and
            duration of throttling and its side effects: head-of-line
            blocking, congestion spreading to unrelated traffic, throttle
            storms, and deadlock.</li>
          <li>Locate the root of congestion. Throttling propagates
            upstream, so several nodes may show large queues. A node whose
            Forward Queue is large but not throttled is a congestion root;
            upstream nodes whose Forward Queues are throttled are victims
            of backpressure. S can then react to the root, e.g., by
            avoiding the root link in load balancing, instead of to the
            victims.</li>
          <li>Estimate queuing delay correctly. While a Forward Queue is
            stopped or slowed by a Throttle Signal, its service rate is
            reduced or zero, so dividing its depth by the link rate
            underestimates its delay (<xref target="source-proc"/>).</li>
        </ul>
        <t>
          HFC reacts hop by hop, within the headroom of a single link.
          FALCON does not replace it: a burst shorter than the control loop
          between S and S_i can still trigger throttling. The purpose of
          this data is to let S reduce sustained buildup that would
          otherwise lead to repeated or prolonged throttling. This benefit
          is greatest where the headroom per link is costly relative to
          the link delay, as on long-distance WAN links.
        </t>
        <t>
          In some devices, flow-control buffer accounting and throttle
          state are kept in the buffer manager and are not directly
          accessible while a packet of a different interface is processed.
          A node that cannot provide a requested value
          <bcp14>SHOULD</bcp14> mark it as unavailable rather than omit
          it.
        </t>
      </section>

      <section anchor="aggregation">
        <name>Per-Hop Collection versus On-Path Aggregation</name>
        <t>
          If transit nodes can perform simple arithmetic in the data
          plane, P' MAY carry aggregate values instead of a per-hop list:
          the running sum of Forward Queue delay (yielding the path
          queuing delay), or the maximum Forward Queue depth or delay
          together with the identifier of the node reporting it (yielding
          the bottleneck), updated by a compare-and-replace operation.
          Where HFC state is collected, P' MAY similarly carry the minimum
          headroom to the Trigger Threshold along the path with the
          identifier of the node reporting it, and a flag set by any node
          that is throttling or being throttled.
          Aggregation keeps the size of P' constant regardless of path
          length.
        </t>
        <t>
          If transit nodes cannot aggregate, P' collects per-hop values,
          and S performs the computation using its knowledge of link
          capacities.
        </t>
      </section>

      <section anchor="source-proc">
        <name>Processing at the Source</name>
        <t>
          The delay of P over the path consists of a baseline component
          (propagation, serialization, and minimum forwarding delay) and a
          queuing component. The queuing component is the sum of the
          Forward Queue delays reported by P'. Where a node reports only
          queue depth, S estimates the queuing delay at S_i as the queue
          depth divided by the service rate of that queue. The service
          rate equals the link rate only when the monitored class is the
          sole active class on the link and is not subject to flow control
          (e.g., HFC throttling, see <xref target="hfc"/>) or shaping; S
          <bcp14>SHOULD</bcp14> prefer node-reported delay estimates when
          available.
        </t>
        <t>
          The baseline component can be obtained by the source as the
          minimum observed RTT (or minimum P-to-P' round trip) over a
          suitable window, divided between directions under an assumption
          of symmetric propagation delay; or, where clocks are
          synchronized, from IOAM timestamps carried in P and P'. The
          estimated current path delay of P is the baseline component
          plus the collected queuing component.
        </t>
        <t>
          S passes the per-hop state, path queuing delay, bottleneck
          information, and congestion indications to its CC or LB
          function. How S reacts is outside the scope of this document.
        </t>
      </section>

      <section anchor="freshness">
        <name>Freshness Analysis</name>
        <t>
          Let d(X,Y) denote the one-way delay from X to Y without
          queuing, and Q_fwd(i) and Q_rev the queuing delays
          encountered after S_i on the forward path and on the
          reverse path from R to S,
          respectively. For the state of S_i:
        </t>
        <ul spacing="normal">
          <li>With forward-direction collection that is echoed back by R,
            the staleness is d(S_i,R) + Q_fwd(i) + d(R,S) + Q_rev. Even
            without queuing, this is at least d(S_i,R) + d(R,S).</li>
          <li>With FALCON, P' traverses the return path in a high-priority
            class, so the staleness is approximately d(S_i,S), which is
            at most d(R,S), or about half the baseline RTT.</li>
        </ul>
        <t>
          The gain is largest for nodes close to the source and grows with
          the amount of forward-path queuing. Since a reaction by S takes
          effect at S_i only after d(S,S_i), the FALCON control loop for a
          bottleneck at S_i is approximately the baseline RTT between S
          and S_i, which is the physical lower bound for any
          source-based reaction.
        </t>
        <t>
          Staleness is distinct from sampling frequency: the interval
          between consecutive P' packets is determined by the sampling
          rate of P (<xref target="sampling"/>). In addition, the
          per-hop samples carried by one P' are taken at slightly
          different times (S_n first, S_1 last), spread over d(S_n,S_1).
        </t>
      </section>
    </section>

    <!-- 6. Implementation and Gap Analysis -->
    <section anchor="implementation">
      <name>Implementation Considerations and Gap Analysis</name>

      <section anchor="encap">
        <name>Encapsulation</name>
        <t>
          In IPv6 networks, P can carry the IOAM trace option in an IPv6
          Hop-by-Hop Options header <xref target="RFC9486"/>. For P', which
          carries both an SRH and an IOAM trace, two approaches are
          possible:
        </t>
        <ul spacing="normal">
          <li>IOAM in a Hop-by-Hop Options header together with the SRH
            <xref target="RFC9486"/>; or</li>
          <li>IOAM data carried in a UDP payload with an SRH flag
            indicating segment-by-segment processing, as proposed in
            <xref target="I-D.song-spring-siam"/>. Since the segment list
            of P' contains a segment for every hop, each transit node is a
            segment endpoint and processes the IOAM data.</li>
        </ul>
      </section>

      <section anchor="gaps">
        <name>Gaps</name>
        <t>
          The following extensions are needed for a standards-based
          realization:
        </t>
        <ol spacing="normal">
          <li>An IOAM data field reporting the state (depth, and optionally
            delay and congestion indication) of the egress queue of the
            interface on which the IOAM packet was received, for a
            specified traffic class.</li>
          <li>An IOAM data field, independent of the specific HFC
            mechanism, reporting for a specified flow-control class the
            Flow-Control Buffer state (occupancy, and optionally headroom
            to the Trigger Threshold and a throttling-upstream indication)
            of the interface through which the IOAM packet leaves the
            node, and a throttled-by-downstream indication for the egress
            queue of the interface on which it was received. The existing
            IOAM queue depth field reports egress queue occupancy, not
            flow-control buffer accounting.</li>
          <li>IOAM data fields or options for on-path aggregation
            (sum, maximum or minimum with location, and flags) as
            described in
            <xref target="aggregation"/>.</li>
          <li>A means for P and P' to indicate the monitored traffic class
            when it differs from the packet's own class.</li>
          <li>A standard encapsulation of IOAM in SRv6 packets, e.g.,
            <xref target="I-D.song-spring-siam"/>.</li>
          <li>An indication in P that requests R to generate P', and the
            parameters for P' (requested data, return traffic class).</li>
          <li>A mechanism to map IOAM node and interface identifiers to
            SRv6 adjacency SIDs, or to record SIDs directly.</li>
        </ol>
      </section>
    </section>

    <!-- 7. Operational Considerations -->
    <section anchor="ops">
      <name>Operational Considerations</name>

      <section anchor="sampling">
        <name>Sampling Rate and Overhead</name>
        <t>
          Generating a P' for every packet is unnecessary and costly at
          high link rates. S <bcp14>SHOULD</bcp14> request recording on a
          configurable subset of packets, e.g., once per configured time
          interval, per configured number of bytes, or per RTT per flow.
          R <bcp14>MUST</bcp14> rate-limit the generation of P', and the
          aggregate volume of P' in the high-priority class
          <bcp14>MUST</bcp14> be bounded so that it cannot starve other
          traffic in that class.
        </t>
        <t>
          The trace data added to P can cause P to exceed the path MTU.
          S <bcp14>SHOULD</bcp14> account for the maximum trace size when
          selecting data packets to carry the instruction, or use
          dedicated probes.
        </t>
        <t>
          The HFC data of <xref target="hfc"/> are useful only where an
          HFC mechanism is enabled for the monitored traffic; S
          <bcp14>SHOULD NOT</bcp14> request them otherwise.
        </t>
      </section>

      <section anchor="partial">
        <name>Partial Deployment</name>
        <t>
          Nodes that do not support FALCON do not record themselves in P
          and are therefore not included in the return segment list; the
          link state between two FALCON nodes separated by a non-capable
          node cannot be observed, and P' traverses the intervening
          segment using normal forwarding, which may not follow the
          reverse of the forward path. S <bcp14>SHOULD</bcp14> detect gaps
          (e.g., via hop-count or TTL information in the trace) and treat
          the corresponding portion of the path as unobserved.
        </t>
      </section>

      <section anchor="path-issues">
        <name>Path Changes, Link Aggregation, and Asymmetry</name>
        <t>
          P' reflects the path taken by the specific P from which it was
          generated. If the monitored flow is re-hashed or re-routed, S
          <bcp14>SHOULD</bcp14> discard state for the old path. Where a
          link is a member of a link aggregation group, the Forward Queue
          is per member link; the recorded interface identifier
          <bcp14>SHOULD</bcp14> identify the member link. FALCON requires
          only that each link be bidirectional; it does not require
          symmetric routing, because P' is explicitly steered. Estimates of
          the baseline delay that assume symmetric propagation are affected
          by asymmetric link delays.
        </t>
      </section>

      <section anchor="applicability">
        <name>Applicability</name>
        <t>
          FALCON is intended for use within a single administrative domain
          that is both an SR domain and an IOAM domain, consistent with
          the deployment scope of <xref target="I-D.song-fann-framework"/>.
          In a data center network the paths are short and the main
          benefit is sub-RTT reaction to incast and transient congestion.
          In a WAN the baseline RTT is large, and the absolute reduction in
          staleness is correspondingly large.
        </t>
      </section>
    </section>

    <!-- 8. Security Considerations -->
    <section anchor="security">
      <name>Security Considerations</name>
      <t>
        The security considerations of IOAM <xref target="RFC9197"/>
        <xref target="RFC9486"/> and SRv6 <xref target="RFC8754"/>
        <xref target="RFC8986"/> apply. In particular, FALCON packets
        <bcp14>MUST</bcp14> be confined to a domain that is both an SR
        domain and an IOAM domain; packets carrying FALCON instructions,
        SRHs, or the P' traffic class <bcp14>MUST</bcp14> be filtered at
        the domain boundary.
      </t>
      <t>
        Additional threats specific to FALCON include:
      </t>
      <ul spacing="normal">
        <li>Forged or modified P' packets could induce S to reduce its
          rate unnecessarily or to move traffic onto a worse path.
          Integrity protection of IOAM data
          <xref target="I-D.ietf-ippm-ioam-data-integrity"/> can mitigate
          modification in transit; S <bcp14>SHOULD</bcp14> accept P' only
          if it corresponds to an outstanding P (e.g., by matching a
          sequence number or nonce carried in P).</li>
        <li>R constructs a segment list from data recorded in P. A crafted
          P could cause R to emit P' along arbitrary paths or toward a
          victim. R <bcp14>MUST</bcp14> verify that all SIDs it places in
          P' belong to the local domain and that the final destination is
          the source address of P, and <bcp14>MUST</bcp14> rate-limit P'
          generation.</li>
        <li>P' uses a high-priority class. If that class is not protected,
          it could be abused to obtain preferential treatment or to
          starve other high-priority traffic. The class
          <bcp14>SHOULD</bcp14> be reserved and policed.</li>
        <li>P' reveals topology, queue occupancy and, where collected,
          buffer thresholds and HFC state to S. Where S is an
          end host that is not fully trusted by the operator, the
          operator <bcp14>SHOULD</bcp14> consider limiting the data
          returned, e.g., to aggregate values only.</li>
      </ul>
    </section>

    <!-- 9. IANA Considerations -->
    <section anchor="iana">
      <name>IANA Considerations</name>
      <t>
        This document has no IANA actions at this time. Future versions
        are expected to request allocations for the extensions listed in
        <xref target="gaps"/>, e.g., new bits in the IOAM Trace-Type
        registry and any flags or code points required for the P'
        request indication.
      </t>
    </section>

  </middle>

  <!-- ============================================================ -->
  <!--  BACK MATTER                                                 -->
  <!-- ============================================================ -->
  <back>
    <references>
      <name>References</name>

      <references>
        <name>Normative References</name>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.2119.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8174.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8754.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8986.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9197.xml"/>
      </references>

      <references>
        <name>Informative References</name>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.2474.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.3168.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9322.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9326.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9486.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9800.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.ietf-fann-problem-statement.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.song-fann-framework.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.song-spring-siam.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.ietf-ippm-ioam-data-integrity.xml"/>
        <xi:include href="https://bib.ietf.org/public/rfc/bibxml3/reference.I-D.han-rtgwg-fine-grained-backpressure.xml"/>

        <reference anchor="IEEE8021Qbb">
          <front>
            <title>IEEE Standard for Local and metropolitan area networks--Media Access Control (MAC) Bridges and Virtual Bridged Local Area Networks--Amendment 17: Priority-based Flow Control</title>
            <author><organization>IEEE</organization></author>
            <date year="2011"/>
          </front>
          <seriesInfo name="IEEE Std" value="802.1Qbb-2011"/>
        </reference>

        <reference anchor="HPCC">
          <front>
            <title>HPCC: High Precision Congestion Control</title>
            <author initials="Y." surname="Li"/>
            <author initials="R." surname="Miao"/>
            <author initials="H." surname="Liu"/>
            <author><organization>et al.</organization></author>
            <date year="2019" month="August"/>
          </front>
          <refcontent>Proceedings of ACM SIGCOMM 2019</refcontent>
        </reference>

        <reference anchor="BOLT">
          <front>
            <title>Bolt: Sub-RTT Congestion Control for Ultra-Low Latency</title>
            <author initials="S." surname="Arslan"/>
            <author initials="Y." surname="Li"/>
            <author initials="G." surname="Kumar"/>
            <author initials="N." surname="Dukkipati"/>
            <date year="2023" month="April"/>
          </front>
          <refcontent>20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23)</refcontent>
        </reference>
      </references>
    </references>

    <!-- Appendix A: Related Work -->
    <section anchor="related">
      <name>Comparison with Related Mechanisms</name>
      <dl newline="true" spacing="normal">
        <dt>Receiver-echoed telemetry (e.g., ECN echo, <xref target="HPCC"/>):</dt>
        <dd>State is sampled on the forward path and returned in
          acknowledgments. Staleness is at least one baseline RTT for the
          first hop and includes forward and reverse queuing.</dd>

        <dt>Switch-originated notifications (e.g., <xref target="BOLT"/>,
          IEEE 802.1Qau QCN, FANN push-based notifications):</dt>
        <dd>The congested node generates a notification directly toward the
          source. Staleness is the same lower bound as FALCON, d(S_i,S),
          but transit nodes must generate packets and must know or derive
          the source to notify. FALCON requires only in-band data
          insertion at transit nodes, and the returned state is scoped to
          the queues used by the source's own traffic, covering the whole
          path in one packet.</dd>

        <dt>IOAM Loopback <xref target="RFC9322"/> and IOAM Direct Export
          <xref target="RFC9326"/>:</dt>
        <dd>Loopback causes transit nodes to return a copy of the packet
          to the encapsulating node, and Direct Export sends data to a
          collector. Both carry state sampled when the forward packet
          passed the node, and neither reads the forward queue on the
          return path.</dd>
      </dl>
    </section>

  </back>
</rfc>
