| Internet-Draft | CATS Service Metrics Operation | October 2026 |
| Zhang, et al. | Expires 13 April 2027 | [Page] |
Computing-Aware Traffic Steering (CATS) optimizes traffic forwarding by considering both computing and networking metrics. The CATS Metrics framework provides a general metric hierarchy and operational guidance for aggregation and normalization. Some deployment scenarios additionally require service-oriented metrics with explicit semantics and units to be consumed directly for joint service-instance and path selection.¶
This document defines an operational approach that does not rely on the Level 1 or Level 2 normalized scores of the CATS Metrics framework. Instead, service sites report service-oriented metrics with explicit semantics and units, such as Global Available Slots and Computing Time, which are directly consumed by the C-PS for joint service-instance and path selection. The document clarifies how these metrics are derived from basic resource information, service reference information, and local policy. It also specifies how the C-PS combines the Computing Service Table with the Network Service Table, and defines update-control and fallback mechanisms suitable for large-scale deployments.¶
Within the unified CATS architecture, this service-oriented operational approach complements the normalized-score approach. The two targeted scenarios are rapid deployment without offline metric-function negotiation and fine-grained, requirement-aware joint selection.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 13 April 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License.¶
The Computing-Aware Traffic Steering (CATS) architecture [RFC10053] aims to steer service traffic to the most suitable service contact instance by evaluating both network state and computing-resource availability. To this end, CATS Service Metric Agents (C-SMAs) collect computing-related metrics, CATS Network Metric Agents (C-NMAs) collect network metrics, and CATS Path Selectors (C-PSes) use the collected information for service contact instance and path selection.¶
The CATS Metrics framework [I-D.ietf-cats-metric-definition-13] defines a multi-level metric framework, including Level 0, Level 1, and Level 2 representations, together with aggregation and normalization functions and operational guidance. Aggregated and normalized metrics provide compact representations that facilitate metric distribution and candidate comparison. However, some deployment scenarios require service-oriented values with explicit semantics and units to support direct, requirement-aware steering decisions.¶
The first targeted scenario is rapid deployment or startup in a large-scale, dynamic, multi-vendor environment. Although the CATS Metrics framework provides procedures for negotiating, synchronizing, and calibrating metric functions, agreeing on and continuously maintaining common functions across heterogeneous equipment and rapidly changing services can be challenging and time-consuming in a large-scale, dynamic, multi-vendor environment. In this scenario, directly reporting service-oriented values such as GAS and Computing Time avoids making offline metric-function negotiation a prerequisite for joint selection.¶
The second targeted scenario is fine-grained, requirement-aware joint selection. A normalized, unitless score in the 0-10 range supports compact candidate ranking, but a value such as 6 does not, by itself, indicate whether a candidate can accommodate 100 concurrent sessions or requests, provide at least 40 Gbps of available bandwidth, keep jitter below 5 ms, or keep total service time below 20 ms. This document therefore preserves explicit computing and network values, including their semantics and units, so that the C-PS can apply such requirements as constraints before jointly selecting a service contact instance and an associated path.¶
Although this document focuses on a single service-provider domain, the explicit semantics and units of these service-oriented metrics may facilitate consistent interpretation and mapping of service information across domain boundaries in future multi-domain deployments.¶
Within the unified CATS architecture, this document defines a service-oriented approach that is complementary to and can operate in parallel with the normalized-score approach of [I-D.ietf-cats-metric-definition-13]. A deployment may use either approach, or operate both as separate control-plane workflows for different service classes. This document defines GAS, Computing Time, and related service information, and specifies how a centralized C-PS combines the Computing Service Table, populated from C-SMA reports, with the Network Service Table to make joint traffic-steering decisions. The metric entries are specified in the companion registry document [I-D.zhangb-cats-service-metric-registry-entries].¶
This document makes use of the terms defined in [RFC10053] and [I-D.ietf-cats-metric-definition-13]. In particular, CS-ID and CSCI-ID are used as CATS identifiers. They provide stable service and service-contact-instance references for lookup and forwarding, but are not treated as computing metrics in this document.¶
Additionally, the following terms are used:¶
Global Available Slots (GAS): The maximum number of concurrent requests/sessions a service site is willing and able to serve for a specific CS-ID at a given time.¶
CS-ID (CATS Service ID): An identifier for a service. It is used as a stable lookup key in the C-PS Computing Service Table.¶
CSCI-ID (CATS Service Contact Instance ID): An identifier for contact information of a service instance that provides a specific CS-ID at a service site. In this document, it is interpreted operationally as a locator, such as an IP address and port number, used to establish the data tunnel.¶
Computing Service Table (CST): A data structure maintained by the C-PS that contains computing-oriented metrics (GAS, Computing Time, etc.) indexed by CS-ID and CSCI-ID. It is populated from C-SMA reports.¶
Network Service Table (NST): A data structure maintained by the C-PS that contains network-oriented metrics (delay, jitter, bandwidth, etc.) indexed by network path identifiers. In an SDN realization, it can be derived from the Traffic Engineering Database (TEDB).¶
Total Service Time (TST): The sum of Computing Time and network delay (Ingress-to-Egress), used as the primary optimization objective in the joint selection algorithm. TST has semantic correspondence to the Level 1 Composed category (end-to-end delay / application-level response time) defined in [I-D.ietf-cats-metric-definition-13].¶
Joint Selection Algorithm: The decision logic used by the C-PS to select the optimal CSCI-ID by simultaneously evaluating both computing and network metrics.¶
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.¶
The CATS working group has made significant progress in defining how computing metrics are collected, distributed, aggregated, and normalized. In particular, the CATS Metrics framework [I-D.ietf-cats-metric-definition-13] defines Level 0, Level 1, and Level 2 metric representations, while [RFC10053] defines the CATS functional architecture and its centralized, distributed, and hybrid deployment models. However, further operational guidance is needed for directly using computing and network information to satisfy explicit service requirements and make joint service contact instance and path selections.¶
Against this background, the following operational gaps arise when a C-PS needs to consume service-oriented information and make fine-grained joint-selection decisions:¶
The Normalization and Coordination Gap: In a real-world multi-vendor network, computing resources are highly heterogeneous. The CATS Metrics framework defines negotiation, synchronization, and calibration procedures for aggregation and normalization functions; however, agreeing on common functions, parameters, and workload profiles, and keeping them synchronized as equipment and services evolve, can introduce significant operational effort.¶
The Actionable-Information Gap: Normalized scores are useful for compact comparison, but a unitless score alone may not preserve all information needed for fine-grained selection. For example, a score of 7 does not, by itself, indicate whether a candidate has sufficient capacity for a specified number of concurrent sessions or requests, satisfies a minimum-bandwidth or maximum-jitter requirement, or meets a maximum total-service-time objective.¶
The Service-Oriented Consumption Gap: Hardware-specific information, including CPU, GPU, or NPU characteristics, can be used locally when deriving service metrics. However, the C-PS primarily needs service-oriented information that answers operational questions such as "Is there sufficient capacity?", "How long is service processing expected to take?", and "Which CSCI-ID can provide the requested service?".¶
The Joint Selection Gap: Even when both computing and network metrics are available, a concrete procedure is needed to combine them into a unified steering decision. [RFC10053] describes a centralized model in which both types of metrics can be collected by logically centralized path-computation logic, but it does not specify how to construct and jointly consume computing and network views, apply explicit service requirements, select a CSCI-ID and associated path, or handle unavailable and stale candidates.¶
To bridge these gaps, this document defines a Service-Oriented Abstraction and a Joint Selection Algorithm. It specifies Mandatory Computing Service Metrics, including Global Available Slots and Computing Time, and Optional Extension Metrics, including Price, Reputation, and Security Level. It also specifies how the C-PS combines the Computing Service Table with the Network Service Table to support executable traffic-steering policies.¶
This operational model targets latency-sensitive, compute-intensive workloads that benefit from explicit service-capacity, processing-time, and network constraints. It complements the general-purpose normalized metric framework and can be deployed independently or alongside that framework for different service classes.¶
This section defines the service information used by CATS control-plane components. Some fields are identifiers or locators, while others are service-oriented metrics. Metric examples follow the structural guidelines specified in Section 4.1 of [I-D.ietf-cats-metric-definition-13]. This document focuses on the semantics and operational use of these metrics, rather than defining new routing, transport, signaling, or wire-encoding mechanisms.¶
[I-D.ietf-cats-metric-definition-13] defines a three-level metric taxonomy: Level 0 metrics, Level 1 category metrics (computing, communication, service, and composed), and a Level 2 normalized metric, together with aggregation and normalization functions and a CATS metric registry. This document does not extend that taxonomy. The service-oriented metrics defined here have the semantic correspondences shown in Table 1, but the operational approach is parallel to the normalized-score approach and does not depend on a Level 1 or Level 2 normalized score.¶
+------------------+---------------------------------------------------+ | Metric | Relationship to the CATS Metric Framework | +------------------+---------------------------------------------------+ | GAS | Semantic correspondence: Level 0 Service-specific | +------------------+---------------------------------------------------+ | Computing Time | Semantic correspondence: Level 0 Service-specific | +------------------+---------------------------------------------------+ | | (registration needs WG consensus) | | Price | Service attribute / policy metric (no level); | +------------------+---------------------------------------------------+ | | (registration needs WG consensus) | | Reputation | Service attribute / policy metric (no level); | +------------------+---------------------------------------------------+ | | (registration needs WG consensus) | | Security Level | Service attribute / policy metric (no level); | +------------------+---------------------------------------------------+
Table 1 expresses semantic correspondence rather than a dependency on the normalized-score pipeline. GAS is consistent with the service-specific metrics anticipated in Section 4.2 of [RFC10053], which notes that computing metrics may include the number of clients accessing a service contact instance at a given time. This document generalizes that operational quantity to the number of concurrent sessions or requests that a service site is willing and able to serve. Computing Time, when combined with network delay (see Section 6.2.3), forms the Total Service Time (TST), which has semantic correspondence to the end-to-end delay or application-level response time described in the Level 1 Composed category of [I-D.ietf-cats-metric-definition-13]. Price, Reputation, and Security Level are operator- or provider-defined service attributes used as constraints or secondary objectives in the Joint Selection Algorithm.¶
The operational approach defined in this document does not rely on any Level 1 or Level 2 normalized score field. GAS and Computing Time are reported and consumed with explicit semantics and units, and the steering decision logic in Section 6.2.3 operates directly on these values without requiring prior negotiation of vendor-specific normalization or aggregation functions. The entries are specified by this document and its companion registry document [I-D.zhangb-cats-service-metric-registry-entries] (see Section 10). Service attributes that are not measured quantities (Price, Reputation, Security Level) are labeled "Service attribute" and are not assigned to a metric level.¶
The current scope is limited to a single service-provider domain, consistent with [RFC10053]. Explicit metric semantics and units may facilitate future multi-domain use; however, inter-domain distribution, trust, policy, and security are outside the scope of this document.¶
The Computing Service Metrics defined in this document are service-oriented metrics. They are estimations produced by each service site based on local monitoring and deployment policy. This allows heterogeneous hardware details and frequent local resource changes to be hidden from the CATS control plane while still exposing actionable information for traffic steering.¶
These metrics can be derived from basic resource metrics, status metrics, service requirements, and local policy at the service site. The public service platform described in [I-D.zhangb-cats-cmas-06] can provide reference information, such as Computing Requirement, Storage Requirement, Reference Computing Time, software dependency, and Reference GAS, that a service site can use when deploying a service.¶
For example, if the resources allocated to a service instance just meet the listed Computing Requirement and Storage Requirement, the service site can use the Reference GAS as a starting value. If more resources are allocated, the reported GAS is evaluated by the service site and is generally expected to be larger than the Reference GAS. Similarly, Computing Time can be measured or estimated based on the runtime behavior of the deployed service instance. The specific derivation algorithm is a local matter and is not standardized by this document.¶
The basic field examples in Section 5.3 and Section 5.4 provide recommended data types, lengths, and units as guidance for protocols that carry these metrics (e.g., BGP-LS extensions or RESTful APIs). This document does not mandate a specific wire format; the exact encoding is a matter for the protocol or transport mechanism used between the C-SMA and the C-PS.¶
To support fair comparison across service sites, the derivation of GAS SHOULD follow a common reference method within an administrative domain; a reference derivation (e.g., the sum over healthy instances of available capacity minus active sessions, with an optional policy safety margin) is provided in [I-D.zhangb-cats-service-metric-registry-entries]. This local derivation consistency is distinct from requiring the Level 1 or Level 2 normalization and aggregation functions discussed in Section 5.4 of [RFC10053].¶
These fields are essential for the C-PS to make fundamental traffic steering decisions. CS-ID and CSCI-ID are identifiers, while GAS and Computing Time are service-oriented metrics.¶
GAS is the core contribution of this metric framework. It represents the maximum number of concurrent requests/sessions that a service site is willing and able to serve for a specific CS-ID through a specific CSCI-ID at a given time.¶
Crucially, GAS acts as a direct abstraction layer over the complex and fluctuating raw computing metrics (CPU, GPU, Memory, Storage) and status metrics (load and health). Instead of exposing highly dynamic raw metrics to the network, the service site absorbs these variations internally. The site can initially provide a GAS value based on its fixed resource allocation and service reference information, and then adjust it according to local policy based on internal status metrics.¶
As the number of concurrent requests increases, the GAS value naturally decreases. Furthermore, the site monitoring system dynamically reduces the GAS value upon detecting abnormal status metrics, such as:¶
Load changes: Sudden increase in internal resources occupied by local users or tasks.¶
Health changes: Sudden performance drop, possibly due to a cyber attack.¶
Reachability: The site crashes or becomes unresponsive.¶
Note: The C-SMA proactively reports significant adjustments to the control plane according to local policy, thresholds, or aggregation intervals. Small per-session changes do not necessarily need to be reported immediately. When GAS drops to 0, it means the instance cannot allocate any more resources, and no new requests will be steered to it. Each reported GAS value is the current estimate at the time of reporting, taken over the configured measurement window.¶
Basic fields:¶
Metric Type: gas Level: Level 0 (Service-specific metrics) Format: unsigned integer Length: four octets Unit: count (concurrent sessions) Statistics: cur Measurement_Window: 10 seconds (default) Value: 500 Source: estimation¶
The time required for the site to perform one service request. The service site can initialize this metric based on service reference information and then measure or estimate it according to the runtime behavior of the deployed service instance. The service site dynamically adjusts this metric based on real-time load and local policy.¶
Computing Time is a critical input to the Joint Selection Algorithm because it represents the processing latency component of the Total Service Time (see Section 6.2.3). It is reported as the mean processing time over the measurement window; implementations MAY additionally report percentile values (e.g., p95) in the Statistics field.¶
Basic fields:¶
Metric Type: comp_time Level: Level 0 (Service-specific metrics) Format: floating point Length: four octets Unit: ms Statistics: mean Measurement_Window: 10 seconds (default) Value: 5 Source: estimation¶
To accommodate advanced traffic-steering scenarios, the following optional fields are defined.¶
Self-defined by the service site to apply administrative or economic billing policies. Price is used as a constraint or secondary objective in the Joint Selection Algorithm (see Section 6.2.4). Price is a service attribute rather than a measured quantity; it is therefore not assigned to a metric level.¶
Basic fields:¶
Metric Type: price Level: Service attribute (no metric level) Format: unsigned integer Length: two octets Unit: unitless (provider-defined pricing unit) Value: 100 Source: nominal¶
A dynamic quality score based on user feedback. Upon completion, if a user experiences long delays or inaccurate results, feedback is returned to the C-PS along with the resource release message. Reputation is used as a tie-breaker or threshold filter in the Joint Selection Algorithm. It is a service attribute rather than a measured quantity and is therefore not assigned to a metric level.¶
The 0-10 range uses the same three score bands as the registered normalized metrics in Section 6 of [I-D.ietf-cats-metric-definition-13]: 0-3 indicates poor reputation (avoid if alternatives exist), 4-7 average (acceptable for steering), and 8-10 excellent (preferred for steering). Reputation remains an optional service attribute in this document, not a Level 1 or Level 2 normalized metric.¶
Basic fields:¶
Metric Type: reputation Level: Service attribute (no metric level) Format: unsigned integer Length: one octet Unit: score (0-10) Value: 8 Source: nominal¶
The Security Level reflects the security status of a service site. A higher score indicates a more secure site. The Security Level is used as a constraint in the Joint Selection Algorithm: a C-PS MUST NOT select a service contact instance with a Security Level below the client's minimum requirement. It is a service attribute rather than a measured quantity and is therefore not assigned to a metric level.¶
Score range: 0-10 (0 indicates the poorest security; 10 indicates optimal security). The range uses the same three score bands as the registered normalized metrics in Section 6 of [I-D.ietf-cats-metric-definition-13]: 0-3 indicates low security posture (not recommended for sensitive traffic), 4-7 medium (acceptable for general traffic), and 8-10 high (preferred for sensitive or regulated traffic). Security Level remains an optional service attribute in this document, not a Level 1 or Level 2 normalized metric.¶
Basic fields:¶
Metric Type: security_level Level: Service attribute (no metric level) Format: unsigned integer Length: one octet Unit: score (0-10) Value: 7 Source: nominal¶
Service sites proactively monitor their internal instances. In large-scale deployments, service sites can use a delta-threshold reporting model. Each service site or C-SMA maintains a local metric cache. Per-session allocation and release events update local GAS values, but do not necessarily trigger immediate reports to the C-PS.¶
Updates are reported when they become operationally significant. Examples include GAS crossing a configured threshold, Computing Time deviating beyond a configured percentage band, or health status changing due to failure, attack detection, or unreachability. A periodic heartbeat or soft-state synchronization can also be used to refresh the C-PS view and avoid stale metrics even when no trigger event occurs.¶
The reporting model defined here aligns with the CATS OAM requirements of [I-D.ietf-cats-oam-fw]: O-REQ-1 (the system SHOULD support both periodic and threshold-triggered reporting) and O-REQ-4 (computing metrics SHOULD be accompanied by timestamps indicating the time of collection so that the C-SMA or C-PS can configure a maximum acceptable staleness threshold). The default update interval is consistent with the "update-interval" parameter of the CATS YANG data model [I-D.ietf-cats-data-model-00].¶
Choosing appropriate protocols for conveying CATS metrics is important. Existing routing protocols such as BGP extensions [RFC4760] and GRASP [RFC8990] may serve as a baseline. For the centralized model in Section 5.3 of [RFC10053], metrics can instead be collected by logically centralized control and path-computation logic. A PCE, an SDN controller and its northbound interfaces [RFC7149] [RFC7426], or another centralized control-plane implementation may realize this model. A systematic analysis of protocol applicability is provided in [I-D.yxl-cats-protocols-applicability]. The specific transport and centralized-control implementation are deployment choices.¶
This section specifies the core operational procedure by which the C-PS combines computing metrics and network metrics to select the optimal CSCI-ID for a service request. This joint decision is the central mechanism that enables Computing-Aware Traffic Steering.¶
The Computing Service Table (CST) is maintained by the C-PS and is populated from C-SMA reports. It contains one entry per (CS-ID, CSCI-ID) pair. Each entry contains the following fields:¶
+-------------------+------------------------------------------------+ | Field | Description | +-------------------+------------------------------------------------+ | CS-ID | The CATS Service Identifier | | CSCI-ID | The Service Contact Instance locator | | | (e.g., IP:Port) | | GAS | Global Available Slots (uint32, count) | | Computing Time | Estimated processing time in ms (float32) | | Price | Service price (optional, uint16) | | Reputation | Quality score 0-10 (optional, uint8) | | Security Level | Security score 0-10 (optional, uint8) | | Last Updated | Timestamp of last C-SMA report | | Expiry Time | Soft-state expiration time | +-------------------+------------------------------------------------+¶
The CST is indexed by CS-ID for fast lookup of all candidate CSCI-IDs for a given service. The C-PS MUST age out entries whose Expiry Time has passed, treating them as unreachable (GAS = 0).¶
The CST is logically separate from the Network Service Table to maintain separation of concerns between computing and network domains.¶
The Network Service Table (NST) is maintained by the C-PS and is populated from C-NMA reports. The NST is a logical data structure maintained by the C-PS, providing a view of network metrics tailored for CATS joint decisions. In an SDN realization, the NST can be derived from the controller's Traffic Engineering Database (TEDB), which aggregates network topology and performance information from the underlay network via protocols such as BGP-LS [RFC8571] or IGP TE extensions [RFC7471] [RFC8570]. Other centralized implementations, including a PCE, can provide an equivalent path-computation view.¶
In the reference architecture, the C-PS accesses the NST rather than the raw TEDB, though implementations MAY co-locate these functions. Instead, a NST Generator component performs the following steps:¶
Filtering: Extracts paths relevant to the CATS domain and the set of Egress CATS-Forwarders.¶
Aggregation: Computes path-level metrics (e.g., one-way delay, available bandwidth) from link-level TEDB data.¶
Indexing: Re-indexes the information by (Ingress CF, Egress CF) pair for fast lookup.¶
The NST MAY be cached and periodically refreshed based on TEDB updates, or MAY be generated on-the-fly based on the current TEDB snapshot. The NST contains one entry per (Ingress, Egress) pair, representing the network path from the client's Ingress CATS-Forwarder to the Egress CATS-Forwarder that connects to the service site. The specific implementation is a local matter.¶
Each entry contains the following fields:¶
+-------------------+------------------------------------------------+ | Field | Description | +-------------------+------------------------------------------------+ | Ingress CF | Ingress CATS-Forwarder identifier | | Egress CF | Egress CATS-Forwarder identifier | | Network Delay | One-way delay in ms (uint32) | | Jitter | Delay variation in ms (uint32, optional) | | Bandwidth | Available bandwidth in Mbps (uint32, optional) | | Loss Rate | Packet loss rate in ppm (uint32, optional) | | Path Attributes | TE attributes (e.g., affinity, color) | | Last Updated | Timestamp of last C-NMA report | +-------------------+------------------------------------------------+¶
The service node may be outside the ingress domain, so this document does not require measuring delay directly from the ingress to the service node. Instead, the Network Delay represents the path from the Ingress CATS-Forwarder to the Egress CATS-Forwarder.¶
This document expresses network delay in milliseconds for consistency with Computing Time in the TST computation (see Section 6.2.3). Network performance metrics such as delay, jitter, and loss are typically distributed via IGP TE metric extensions for OSPF [RFC7471] and IS-IS [RFC8570], or via BGP-LS TE performance metric extensions [RFC8571], which use microseconds for delay values. When consuming such advertisements, the C-PS SHOULD convert the received values to a common time base (e.g., milliseconds) before populating the NST and computing TST, to avoid unit mismatches.¶
The C-PS MUST ensure that the NST and CST are synchronized in time: when evaluating a candidate, the C-PS SHOULD use NST entries whose Last Updated timestamp is within a reasonable window of the CST entry's Last Updated timestamp (e.g., within 30 seconds) to avoid decisions based on stale combinations.¶
Network performance metrics such as delay, jitter, and loss can be distributed via IGP TE metric extensions for OSPF [RFC7471] and IS-IS [RFC8570], or via BGP-LS TE performance metric extensions [RFC8571]. The C-PS can consume these standard protocol advertisements to populate the NST.¶
The algorithm in this section applies to the centralized CATS model described in Section 5.3 of [RFC10053]. For an incoming service request, logically centralized path-computation logic, such as a PCE, processes the requirements associated with that request and combines computing-related metrics from C-SMAs with network metrics from C-NMAs to compute candidate paths to service contact instances. Based on these metrics and paths, the centralized C-PS jointly selects the appropriate CSCI-ID and associated path and synchronizes the result with the relevant C-TC. Procedures for distributed or hybrid selection, or for precomputing paths independently of an incoming service request, are outside the scope of this algorithm.¶
When a service request received by an Ingress CATS-Forwarder is presented to the centralized C-PS, the C-PS executes the following Joint Selection Algorithm to determine the selected CSCI-ID and associated path:¶
The C-PS queries the CST using the requested CS-ID. It retrieves all entries matching the CS-ID. Entries with GAS = 0 or expired entries are discarded. This produces the Candidate Set:¶
Candidate_Set = { (CSCI-ID_i, GAS_i, CompTime_i, ...) |
CST[CS-ID, CSCI-ID_i] exists AND
CST.GAS > 0 AND
CST.Expiry > now }
¶
The C-PS applies hard constraints to eliminate infeasible candidates:¶
Security Level: If the client or policy specifies a minimum security label S_min, remove all candidates where Security_Level < S_min.¶
Reputation: If the client or policy specifies a minimum reputation R_min, remove all candidates where Reputation < R_min.¶
Price: If the client or policy specifies a maximum price M_max, remove all candidates where Price > M_max.¶
GAS Sufficiency: If local policy indicates a required session count N_req, remove all candidates where GAS < N_req. (Note: typically N_req = 1 for a single session.)¶
After constraint filtering, if the Candidate_Set is empty, the C-PS proceeds to Fallback (Section 6.2.5).¶
For each remaining candidate CSCI-ID_i, the C-PS determines the corresponding Egress CATS-Forwarder Egress_i (based on the CSCI-ID's network attachment point). It then queries the NST for the path from the client's Ingress CATS-Forwarder to Egress_i:¶
Path_i = NST[Ingress, Egress_i]
¶
If no NST entry exists for a candidate, that candidate is removed from the Candidate_Set.¶
For each remaining candidate, the C-PS computes the Total Service Time (TST):¶
TST_i = Computing_Time_i + Network_Delay_i
¶
Where:¶
Computing_Time_i is from the CST entry for CSCI-ID_i.¶
Network_Delay_i is from the NST entry for Path_i.¶
TST has semantic correspondence to the Level 1 Composed category (end-to-end delay) of [I-D.ietf-cats-metric-definition-13], and corresponds to the "Joint Performance Measurement" (total latency = network transmission time + service processing time) defined in the CATS OAM framework [I-D.ietf-cats-oam-fw]. Computing_Time_i and Network_Delay_i MUST be expressed on the same time base (see Section 6.2.2).¶
The C-PS applies network-side hard constraints to eliminate infeasible candidates:¶
Bandwidth: If the client or policy specifies a minimum bandwidth BW_min, remove all candidates where Bandwidth < BW_min.¶
Jitter: If the client or policy specifies a maximum jitter J_max, remove all candidates where Jitter > J_max.¶
Loss Rate: If the client or policy specifies a maximum loss rate LR_max, remove all candidates where Loss_Rate > LR_max.¶
Latency: If the client or policy specifies a maximum total latency L_max, remove all candidates where TST_i > L_max.¶
After constraint filtering, if the Candidate_Set is empty, the C-PS proceeds to Fallback (Section 6.2.5).¶
If multiple candidates remain after all hard constraints have been applied, the C-PS performs optimization selection. The client service requirement MAY explicitly designate one or more metrics as optimization objectives (i.e., find the best value for that metric), as opposed to constraints. Optimization objectives MUST be expressed as an explicit list in the client service requirement; they MUST NOT be inferred from a reserved metric value such as 0, because values such as Price = 0 (free service) or Security_Level = 0 (poorest security) are legitimate metric values. The C-PS processes optimization objectives in the following priority order:¶
Optimize Security Level: Optimal_CSCI-ID = argmax_{i in Candidate_Set} (Security_Level_i)¶
Optimize Reputation: Optimal_CSCI-ID = argmax_{i in Candidate_Set} (Reputation_i)¶
Optimize Price (minimize): Optimal_CSCI-ID = argmin_{i in Candidate_Set} (Price_i)¶
Optimize GAS: Optimal_CSCI-ID = argmax_{i in Candidate_Set} (GAS_i)¶
Optimize Bandwidth: Optimal_CSCI-ID = argmax_{i in Candidate_Set} (Bandwidth_i)¶
Optimize Jitter (minimize): Optimal_CSCI-ID = argmin_{i in Candidate_Set} (Jitter_i)¶
Optimize Loss Rate (minimize): Optimal_CSCI-ID = argmin_{i in Candidate_Set} (Loss_Rate_i)¶
Optimize Latency (minimize): Optimal_CSCI-ID = argmin_{i in Candidate_Set} (TST_i)¶
If no optimization objectives are specified, the C-PS applies soft constraints and preference rules listed in the client service requirement. The processing of soft constraints follows the same logic as hard constraints, but violation of a soft constraint does not eliminate a candidate; instead, it contributes to a penalty in the candidate's score.¶
If still multiple candidates exist after optimization and soft-constraint processing, the C-PS applies tie-breaking rules in the following priority:¶
The following figure illustrates the data flow and decision logic:¶
+-------------------------+ +-------------------------+
| Computing Service Table | | Network Service Table |
| (CS-ID, CSCI-ID, GAS, | | (Ingress, Egress, |
| Comp Time, Price, etc.) | | Delay, Jitter, BW) |
+-----------+-------------+ +-----------+-------------+
| |
| 1. Lookup by CS-ID | 3. Lookup by
| (filter GAS>0) | (Ingress, Egress)
v v
+----------------------------------------------------------+
| C-PS Joint Selector |
| |
| 2. Apply Computing Hard Constraints |
| (Security, Reputation, Price, GAS) |
| |
| 4. Compute TST = Comp Time + Network Delay |
| |
| 5. Apply Network Hard Constraints |
| (Bandwidth, Jitter, Loss, Latency) |
| |
| 6. Optimize / Tie-break |
| |
| 7. Return CSCI-ID to Ingress CATS-Forwarder |
+----------------------------------------------------------+
|
v
+----------+----------+
| Ingress CATS-FW |
| (Encapsulate & |
| Forward) |
+----------+----------+
|
v
+----------+----------+
| Egress CATS-FW |
| (Decapsulate & |
| Deliver to SCI) |
+---------------------+
The basic Joint Selection Algorithm executes according to the client's service requirement parameters. However, production deployments often require multi-objective optimization that balances competing goals. This section defines extensions to the basic algorithm.¶
The C-PS MAY use a weighted objective function that combines multiple metrics:¶
Score_i = w1 * n(TST_i) + w2 * n(Price_i) + w3 * (1 - n(GAS_i))
+ w4 * (1 - n(Reputation_i)) + w5 * (1 - n(Security_Level_i))
+ w6 * (1 - n(Bandwidth_i)) + w7 * n(Loss_Rate_i)
+ w8 * n(Jitter_i)
¶
Where w1, w2, w3, w4, w5, w6, w7, w8 are non-negative weights configured by the manager, and n(x) denotes min-max normalization of each metric to the range [0,1] based on implementation-configured minimum and maximum expected values per metric (when max = min, n(x) = 0.5). Terms for metrics that are maximized (GAS, Reputation, Security Level, Bandwidth) are expressed as (1 - n(value)) so that a better value contributes a lower score; reciprocal terms such as 1/GAS are avoided because zero is a legitimate value for some metrics (e.g., Security Level = 0, Price = 0) and would cause division by zero. Because all terms are unitless and bounded, weights can be compared and configured meaningfully. The C-PS selects the candidate with the minimum Score_i. By default, w1 = 1 and w2 = w3 = w4 = w5 = w6 = w7 = w8 = 0 (pure TST minimization). Weights are normalized such that each term contributes proportionally to its configured priority.¶
Instead of (or in addition to) optimization, the C-PS MAY apply constraint-based selection:¶
Delay Budget: TST_i <= TST_max. Remove candidates exceeding the maximum acceptable total service time.¶
Price Budget: Price_i <= Price_max. Remove candidates exceeding the maximum acceptable price.¶
Bandwidth Requirement: Bandwidth_i >= BW_min. Remove candidates that cannot provide sufficient network bandwidth.¶
Affinity Requirement: If the client requires session affinity to a previously selected CSCI-ID, and that CSCI-ID is still in the Candidate_Set, the C-PS MAY bypass the optimization and select the affined CSCI-ID directly.¶
For scalability, the C-PS MAY perform hierarchical selection:¶
First, select the best Egress CATS-Forwarder based on network metrics alone (e.g., minimum Network Delay).¶
Then, among the SCIs reachable via that Egress CATS-Forwarder, select the best CSCI-ID based on computing metrics.¶
This reduces the search space and simplifies the decision, but may miss globally optimal solutions where a slightly longer network path leads to a significantly better computing site.¶
The C-PS MAY support policy-driven overrides that take precedence over the optimization algorithm:¶
Geo-fencing: Always prefer service sites within a specific geographic region.¶
Provider preference: Always prefer a specific service provider.¶
Maintenance avoidance: Avoid service sites under maintenance.¶
Load balancing: Distribute traffic evenly across multiple sites even if one has slightly better TST.¶
When the Joint Selection Algorithm cannot produce a valid candidate, the C-PS MUST execute fallback procedures:¶
If the Candidate_Set is empty after applying constraints, the C-PS MAY apply some implementation-specific relaxation policies, for example:¶
Increase the acceptable Price threshold by 20%.¶
Decrease the acceptable Security Level by 1 point.¶
Decrease the minimum Reputation threshold by 1 point.¶
Decrease the minimum GAS threshold by 1 point.¶
Decrease the acceptable Bandwidth threshold by 20%.¶
Increase the acceptable Jitter threshold by 20%.¶
Increase the acceptable Loss Rate threshold by 20%.¶
Increase the acceptable Latency threshold by 20%.¶
Accept DEGRADED service sites (if previously excluded).¶
Re-run the Joint Selection Algorithm with relaxed constraints.¶
If the selected CSCI-ID becomes unreachable after the session is established (detected by C-PS via C-SMA withdrawal or path failure), the C-PS:¶
This section provides detailed use cases that illustrate the integrated routing logic of the Joint Selection Algorithm in the centralized CATS model. The scenarios demonstrate how the centralized C-PS combines the Computing Service Table and the Network Service Table under different operational conditions. In these examples, ARS, VRS, and LLMS denote an Augmented Reality service, a Virtual Reality service, and a Large Language Model inference service, respectively; they are illustrative service names used as CS-IDs, not router models.¶
Consider a CATS domain with three service sites providing computing services. The topology is as follows:¶
Service Site 2 (SS2) Service Site 3 (SS3)
+------------------+ +------------------+
| SCI-1: ARS | | SCI-3: ARS |
| 188.3.67.3:67 | | 188.3.67.4:69 |
| GAS=400, CT=5 | | GAS=600, CT=6 |
| Price=10 | | Price=5 |
| | | |
| SCI-2: VRS | | SCI-4: LLMS |
| 188.3.67.3:68 | | 188.3.67.4:70 |
| GAS=100, CT=15 | | GAS=300, CT=12 |
| Price=20 | | Price=15 |
+--------+---------+ +--------+---------+
| |
Egress CF-2 Egress CF-3
(188.3.67.3) (188.3.67.4)
^ ^
| |
+--------------+-----------------------------------+-------------+
| Underlay Network |
| (P-routers) |
+--------------+-----------------------------------+-------------+
^ ^
| |
+-----+------+ +-----+------+
| Ingress | | C-NMA |
| CATS-FW 1 | +-----+------+
| 10.0.0.1 | |
+-----+------+ |
^ |
| Selected CSCI-ID and path | Network metrics
| |
+--------+-----------------------------------+--------+
| Centralized C-PS / Path-Computation Logic |
| (CST + NST) |
+--------------------------+--------------------------+
^
| Service metrics
+----+----+
| C-SMAs |
+---------+
¶
The C-SMAs at each service site report their local service information to the logically centralized C-PS and associated path-computation logic, forming the Computing Service Table (CST). The C-NMA reports network information to the same centralized logic, forming the Network Service Table (NST). After joint selection, the centralized C-PS supplies the selected CSCI-ID and associated path or steering instructions to the Ingress CATS-Forwarder. The figure shows logical functional placement and does not require the functions to be implemented on particular physical devices.¶
Computing Service Table (CST):¶
+=======+===================+=====+===============+===================+ | CS-ID | CSCI-ID (IP:Port) | GAS | Comp Time(ms) | Price (Optional) | +=======+===================+=====+===============+===================+ | ARS | 188.3.67.3:67 | 400 | 5 | 10 | +-------+-------------------+-----+---------------+-------------------+ | VRS | 188.3.67.3:68 | 100 | 15 | 20 | +-------+-------------------+-----+---------------+-------------------+ | ARS | 188.3.67.4:69 | 600 | 6 | 5 | +-------+-------------------+-----+---------------+-------------------+ | LLMS | 188.3.67.4:70 | 300 | 12 | 15 | +-------+-------------------+-----+---------------+-------------------+¶
Network Service Table (NST):¶
+========================+=======================+============+ | Ingress CATS-Forwarder | Egress CATS-Forwarder | Network | | | | Delay (ms) | +========================+=======================+============+ | 10.0.0.1 | 188.3.67.3 (CF-2) | 8 | +------------------------+-----------------------+------------+ | 10.0.0.1 | 188.3.67.4 (CF-3) | 6 | +------------------------+-----------------------+------------+¶
Note: In this example, the Egress CATS-Forwarder identifiers are represented by their attachment IP addresses for simplicity. In a real deployment, they would be router IDs or loopback addresses.¶
A client at Ingress CATS-Forwarder 1 (10.0.0.1) requests the ARS service with the requirement for the shortest Total Service Time.¶
The C-PS queries the CST for CS-ID = "ARS". Candidates:¶
CSCI-ID 188.3.67.3:67 (SS2, GAS=400, CT=5ms, Price=10)¶
CSCI-ID 188.3.67.4:69 (SS3, GAS=600, CT=6ms, Price=5)¶
Both have GAS > 0 and are not expired.¶
The C-PS compares TST values:¶
The C-PS selects 188.3.67.4:69 (Service Site 3) because it offers the shortest Total Service Time, even though its computing time (6 ms) is slightly higher than candidate 188.3.67.3:67 (5 ms). The shorter network path (6 ms vs. 8 ms) compensates for the slightly higher computing time.¶
This scenario demonstrates that CATS selects the globally optimal combination of computing and network performance, not just the best computing site or the best network path in isolation.¶
Suppose the network path to Service Site 3 experiences congestion, and the C-NMA updates the NST:¶
Updated Network Service Table (NST):¶
+========================+=======================+============+ | Ingress CATS-Forwarder | Egress CATS-Forwarder | Network | | | | Delay (ms) | +========================+=======================+============+ | 10.0.0.1 | 188.3.67.3 (CF-2) | 8 | +------------------------+-----------------------+------------+ | 10.0.0.1 | 188.3.67.4 (CF-3) | 20 | <-- changed +------------------------+-----------------------+------------+¶
The C-PS re-evaluates the ARS request:¶
The C-PS dynamically re-selects 188.3.67.3:67 (Service Site 2) and updates the forwarding rules at Ingress CATS-Forwarder 1. New client requests for ARS are steered to Service Site 2 until the network congestion to Service Site 3 subsides.¶
This scenario demonstrates CATS's ability to dynamically adapt to network condition changes and re-steer traffic to maintain optimal Total Service Time.¶
Suppose Service Site 3 experiences a partial failure: SCI-3 (ARS) becomes DEGRADED, and its GAS drops from 600 to 50. The C-SMA at Service Site 3 reports the updated metrics to the C-PS:¶
Updated Computing Service Table (CST):¶
+=======+===================+=====+===============+===================+ | CS-ID | CSCI-ID (IP:Port) | GAS | Comp Time(ms) | Price (Optional) | +=======+===================+=====+===============+===================+ | ARS | 188.3.67.3:67 | 400 | 5 | 10 | +-------+-------------------+-----+---------------+-------------------+ | VRS | 188.3.67.3:68 | 100 | 15 | 20 | +-------+-------------------+-----+---------------+-------------------+ | ARS | 188.3.67.4:69 | 50 | 25 | 5 | <-- changed +-------+-------------------+-----+---------------+-------------------+ | LLMS | 188.3.67.4:70 | 300 | 12 | 15 | +-------+-------------------+-----+---------------+-------------------+¶
Note: Computing Time increased to 25 ms due to degraded performance.¶
The C-PS re-evaluates the ARS request:¶
Now suppose the failure worsens: SCI-3 becomes UNHEALTHY, and GAS drops to 0. The C-SMA reports the withdrawal:¶
Updated Computing Service Table (CST):¶
+=======+===================+=====+===============+===================+ | CS-ID | CSCI-ID (IP:Port) | GAS | Comp Time(ms) | Price (Optional) | +=======+===================+=====+===============+===================+ | ARS | 188.3.67.3:67 | 400 | 5 | 10 | +-------+-------------------+-----+---------------+-------------------+ | VRS | 188.3.67.3:68 | 100 | 15 | 20 | +-------+-------------------+-----+---------------+-------------------+ | ARS | 188.3.67.4:69 | 0 | 25 | 5 | <-- GAS=0 +-------+-------------------+-----+---------------+-------------------+ | LLMS | 188.3.67.4:70 | 300 | 12 | 15 | +-------+-------------------+-----+---------------+-------------------+¶
The C-PS selects 188.3.67.3:67 as the only available candidate. This scenario demonstrates automatic failover when a service site becomes unavailable.¶
If ALL ARS candidates have GAS = 0 (e.g., both sites fail), the C-PS triggers Fallback Level 2 (Section 6.2.5): it falls back to pure network-based steering or returns a "Service Unavailable" indication to the client.¶
Consider a client request for ARS with the following policy:¶
Primary objective: Minimize Total Service Time.¶
Hard constraint: Price <= 8.¶
Secondary objective: Maximize GAS (prefer sites with more available capacity).¶
The CST and NST are as in the initial state (Scenario A):¶
Computing Service Table (CST):¶
+=======+===================+=====+===============+===================+ | CS-ID | CSCI-ID (IP:Port) | GAS | Comp Time(ms) | Price (Optional) | +=======+===================+=====+===============+===================+ | ARS | 188.3.67.3:67 | 400 | 5 | 10 | +-------+-------------------+-----+---------------+-------------------+ | ARS | 188.3.67.4:69 | 600 | 6 | 5 | +-------+-------------------+-----+---------------+-------------------+¶
Network Service Table (NST):¶
+========================+=======================+============+ | Ingress CATS-Forwarder | Egress CATS-Forwarder | Network | | | | Delay (ms) | +========================+=======================+============+ | 10.0.0.1 | 188.3.67.3 (CF-2) | 8 | +------------------------+-----------------------+------------+ | 10.0.0.1 | 188.3.67.4 (CF-3) | 6 | +------------------------+-----------------------+------------+¶
Apply Price <= 8:¶
Candidate_Set = { 188.3.67.4:69 }.¶
This scenario demonstrates how hard constraints (Price <= 8) filter out expensive candidates before optimization. Even though 188.3.67.3:67 has a slightly better computing time (5 ms vs. 6 ms), it is excluded due to price.¶
Now suppose the client policy changes to Price <= 12 (relaxing the constraint). Both candidates are eligible:¶
However, if the policy includes a secondary objective (maximize GAS) and both candidates have the same TST (e.g., due to a network change making both paths equal), the tie-breaker rule selects the candidate with higher GAS. If 188.3.67.4:69 has GAS = 600 and 188.3.67.3:67 has GAS = 400, the C-PS prefers 188.3.67.4:69.¶
This scenario demonstrates the interaction between hard constraints, primary optimization (TST), and secondary tie-breaking (GAS).¶
In large-scale deployments with hundreds of service sites and thousands of service instances, unconstrained metric updates can overwhelm the C-PS and the control-plane network. This section defines update control mechanisms. These mechanisms are consistent with the scalability considerations of [I-D.ietf-cats-metric-definition-13] (Section 3.1) and with the CATS OAM framework [I-D.ietf-cats-oam-fw] (O-REQ-1: balanced periodic and threshold-triggered reporting; O-REQ-4: freshness handling).¶
Each C-SMA maintains a local metric cache. Per-session GAS changes are accumulated locally. The C-SMA reports an update only when:¶
GAS changes by more than a configured absolute threshold (e.g., 50 slots) or relative threshold (e.g., 10%).¶
Computing Time changes by more than a configured relative threshold (e.g., 20%).¶
Health status changes (HEALTHY <-> DEGRADED <-> UNHEALTHY).¶
A periodic heartbeat interval expires (e.g., every 60 seconds).¶
This section complements the general security requirements for CATS metrics defined in Section 8 of [I-D.ietf-cats-metric-definition-13]: integrity (SEC-1), authenticity (SEC-2), controllability (SEC-3), freshness (SEC-4), and confidentiality (SEC-5). The requirements below are specific to the Computing Service Metrics and the Joint Selection Algorithm defined in this document; they MUST be applied in addition to, and consistent with, the general requirements.¶
The dynamic reporting of Service Metrics introduces potential attack vectors. Authentication mechanisms between service sites and C-SMAs MUST be enforced. The Security Level (Section 5.4.3) can be utilized by the C-PS to prevent routing sensitive traffic to compromised sites.¶
Service Metric reports influence service selection and therefore need integrity protection, source authentication, and authorization checks. Deployments should protect against forged, replayed, or stale metric reports, for example by using freshness information and aging out old metric state. Implementations should also consider rate limiting or aggregation policies so that abnormal local events do not create excessive update bursts toward the C-PS.¶
The Joint Selection Algorithm relies on the integrity of both the CST and the NST. If an attacker injects false network delays into the NST or false computing times into the CST, traffic may be misdirected to suboptimal or compromised service sites. Therefore, both the C-NMA and C-SMA MUST authenticate their reports to the C-PS, and the C-PS MUST validate the freshness and plausibility of all inputs.¶
Specific threats related to individual Computing Service Metrics include:¶
Finally, the C-PS itself is a critical security component. If compromised, it could steer all traffic to an attacker-controlled service site. The C-PS SHOULD run in a hardened environment, and its policy configuration SHOULD be protected against unauthorized modification.¶
This section follows the guidelines in [RFC8126].¶
[I-D.ietf-cats-metric-definition-13] requests IANA to create the "CATS Metrics" registry under the "Computing-Aware Traffic Steering (CATS)" heading. To keep a single point of registration for CATS metrics, this document does not request a separate registry. Instead, the service-oriented metric entries defined in this document (GAS, Computing Time, Price, Reputation, and Security Level) are registered as entries in the "CATS Metrics" registry, in a dedicated "Computing Service Metrics" section of that registry.¶
Each registry entry contains the following fields:¶
Identifier: A unique integer assigned by IANA.¶
Name: The formal metric name.¶
URI: A stable URI reference for the metric entry.¶
Description: A brief description of the metric.¶
Change Controller: The entity responsible for the metric definition (IETF for entries defined in this document and its companion registry document).¶
Version: The version of the metric definition.¶
The registration policy for these entries is "IETF Review" [RFC8126], consistent with the CATS metric framework, because the entries influence traffic steering decisions. The policy is aligned between this document and the companion registry document [I-D.zhangb-cats-service-metric-registry-entries].¶
The initial entries are listed below. Full registry definitions (including Summary, Metric Definition, Method of Measurement, Output, and Administrative Items) are provided in the companion document [I-D.zhangb-cats-service-metric-registry-entries], which serves as the authoritative specification for the initial entries. The naming convention follows the pattern defined in that document ("Svc_" prefix), extending the "Norm_"/"Comb_" naming of [I-D.ietf-cats-metric-definition-13] to service-oriented metrics that are not normalized scores (e.g., GAS) or that are score attributes (e.g., Reputation).¶
+------------+---------------------------------------+
| Identifier | Metric Name (short) |
+------------+---------------------------------------+
| TBD1 | GAS |
| TBD2 | Computing Time |
| TBD3 | Price |
| TBD4 | Reputation |
| TBD5 | Security Level |
+------------+---------------------------------------+
¶
Note: The allocation of identifiers and the exact relationship between this registry section and the "CATS Metrics" registry are to be finalized in coordination with the CATS working group before publication.¶