<?xml version="1.0" encoding="utf-8"?>
<!--
     draft-rfcxml-general-template-standard-00

     This template includes examples of the most commonly used features of RFCXML with comments
     explaining how to customise them. This template can be quickly turned into an I-D by editing
     the examples provided. Look for [REPLACE], [REPLACE/DELETE], [CHECK] and edit accordingly.
     Note - 'DELETE' means delete the element or attribute, not just the contents.

     Documentation is at https://authors.ietf.org/en/templates-and-schemas
-->
<?xml-model href="rfc7991bis.rnc"?>  <!-- Required for schema validation and schema-aware editing -->
<!-- <?xml-stylesheet type="text/xsl" href="rfc2629.xslt" ?> -->
<!-- This third-party XSLT can be enabled for direct transformations in XML processors, including most browsers -->


<!DOCTYPE rfc [
  <!ENTITY nbsp    "&#160;">
  <!ENTITY zwsp   "&#8203;">
  <!ENTITY nbhy   "&#8209;">
  <!ENTITY wj     "&#8288;">
]>
<!-- If further character entities are required then they should be added to the DOCTYPE above.
     Use of an external entity file is not recommended. -->

<rfc
  xmlns:xi="http://www.w3.org/2001/XInclude"
  category="std"
  docName="draft-ietf-mlcodec-opus-speech-coding-enhancement-04"
  ipr="trust200902"
  obsoletes=""
  updates="6716"
  submissionType="IETF"
  xml:lang="en"
  version="3">
<!-- [REPLACE]
       * docName with name of your draft
     [CHECK]
       * category should be one of std, bcp, info, exp, historic
       * ipr should be one of trust200902, noModificationTrust200902, noDerivativesTrust200902, pre5378Trust200902
       * updates can be an RFC number as NNNN
       * obsoletes can be an RFC number as NNNN
-->

  <front>
    <title abbrev="Opus Speech Coding Enhancement">Integration of Speech Codec Enhancement Algorithms into the Opus Codec</title>
    <!--  [REPLACE/DELETE] abbrev. The abbreviated title is required if the full title is longer than 39 characters -->

    <seriesInfo name="Internet-Draft" value="draft-ietf-mlcodec-opus-speech-coding-enhancement-04"/>

    <author fullname="Jan" initials="J." role="editor" surname="Buethe">
      <organization>Meta Platforms Inc.</organization>
      <address>
        <postal>
          <country>US</country>
          <!-- Uses two letter country code -->
        </postal>
        <email>jan.buethe@googlemail.com</email>
      </address>
    </author>

    <author fullname="Jean-Marc" initials="J.-M." surname="Valin">
      <organization>Google</organization>
      <address>
        <postal>
          <country>CA</country>
          <!-- Uses two letter country code -->
        </postal>
        <email>jmvalin@jmvalin.ca</email>
      </address>
    </author>

    <date year="2026" month="July"/>
    <!-- On draft subbmission:
         * If only the current year is specified, the current day and month will be used.
         * If the month and year are both specified and are the current ones, the current day will
           be used
         * If the year is not the current one, it is necessary to specify at least a month and day="1" will be used.
    -->

    <area>Applications and Real-Time</area>
    <workgroup>Machine Learning for Audio Coding</workgroup>
    <!-- "Internet Engineering Task Force" is fine for individual submissions.  If this element is
          not present, the default is "Network Working Group", which is used by the RFC Editor as
          a nod to the history of the RFC Series. -->

    <keyword>Opus, RFC6716</keyword>
    <!-- [REPLACE/DELETE]. Multiple allowed.  Keywords are incorporated into HTML output files for
         use by search engines. -->

    <abstract>
      <t>This document proposes a set of requirements for integrating a speech codec enhancement algorithm into the Opus codec <xref target="RFC6716"/>.</t>
    </abstract>
  </front>

  <middle>
    <section>
      <name>Introduction</name>
      <t>
        Since the specification of the original Opus codec <xref target="RFC6716"/>, new data-driven speech codec enhancement algorithms have emerged that outperform classical
        enhancement algorithms by a large margin. Using such enhancement algorithms to improve the quality of the Opus speech codec SILK requires an update of <xref target="RFC6716"/>
        since SILK is an embedded coding mode and changing the output of the SILK decoder will lead to a violation of the Opus conformance criteria. The purpose of this document
        is hence to update <xref target="RFC6716"/> to enable the use of a speech codec enhancement algorithm. Specifically, this document defines the notion of a SILK enhancement
        algorithm and sets forth a list of requirements, some mandatory, some optional, that aim to ensure
      </t>
      <ol type="(%d)">
        <li>consistent performance of the enhancement algorithm itself, </li>
        <li>preservation of decoder performance (e.g. seamless mode switching), and</li>
        <li>preservation of basic interoperability when tuning the Opus encoder for use with an enhanced decoder.</li>
      </ol>
      <t>
        While the first two objectives target the Opus decoder alone, the third objective introduces new restrictions on the Opus encoder. However, these are not expected
        to interfere with any existing implementation of an Opus encoder since they target potential interoperability issues arising from new incentives connected to
        the possibility to enhance the Opus decoder.
      </t>
      <t>
        The approach of specifying requirements instead of specifying the enhancement algorithm itself has the advantage of allowing the Opus decoder to benefit from
        future improvements in a field that currently sees rapid development. Still, a description of the linear-adaptive coding enhancer (LACE) and its integration
        into the Opus decoder is included as an illustrative example for a SILK enhancement algorithm.
      </t>

      <section>
        <name>Requirements Language</name>
        <t>The key words "MUST", "MUST NOT", "REQUIRED", "SHALL",
          "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT
          RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be
          interpreted as described in BCP 14 <xref target="RFC2119"/>
          <xref target="RFC8174"/> when, and only when, they appear in
          all capitals, as shown here.</t>
      </section>
      <!-- [CHECK] The 'Requirements Language' section is optional -->

    </section>

    <section anchor="LACE">
      <name>An Illustrative Example</name>
      <t>
        We use the linear-adaptive coding enhancer (LACE) <xref target="lace-paper"/> as an illustrative example to highlight the specific challenges of integrating a speech codec enhancement algorithm into the Opus decoder.
         LACE is trained to enhance the output signal of the SILK decoder, the speech coding mode of Opus, and <xref target="opus-with-lace"/> depicts a high-level overview of the Opus decoder with LACE added as
         an enhancement algorithm.
      </t>
      <t>
        The first requirement for a speech coding enhancement algorithm concerns the performance of the algorithm itself. In this example it relates to the question of how the SILK decoder output compares to the LACE output.
        In <xref target="lace-paper"/> this has been evaluated on clean speech samples using a P.808 listening test <xref target="p.808"/> as well as the objective method PESQ, which showed consistent improvement for all tested bitrates.
        For a general enhancement algorithm it will be necessary to specify testing material and performance criteria to prevent unintended quality degradation of the Opus codec.
      </t>
      <t>
        The second requirement concerns performance of the Opus decoder as a whole. Depending on the bitstream the decoder may have to perform mode switching, e.g. between SILK and CELT, or it may combine the SILK and CELT outputs
        when the codec operates in hybrid mode. Changes to the SILK output signal by an enhancement algorithm, such as added delay, phase shifts, or level alterations can therefore negatively impact the performance
        of the Opus decoder even if the first requirement is met. LACE solves this problem by adding no delay and by being approximately phase and level preserving. However, since many enhancement algorithms are non-causal and
        non-phase-preserving, these requirements may be too strict for a general enhancement algorithm.
      </t>

      <t>
        The third requirement concerns interoperability. The Opus specification provides significant freedom for tuning the encoder and the presence of an enhancement algorithm in the decoder may change the optimal encoding choices
        significantly. In the present example encoding e.g. wideband content at 6 kb/s still leads to fair-to-good quality when using the LACE-enhanced decoder while the quality of a legacy decoder is significantly worse.
        To make full use of these new enhancement algorithms, such encoder tunings should be allowed but basic interoperability with legacy decoders or other enhanced decoders needs to be ensured.
      </t>

      <figure anchor="opus-with-lace">
        <name>A simplified Opus decoder diagram including LACE as enhancement module</name>
        <artwork type="ascii-art" name="lace.txt">
      <![CDATA[
                 ┌──────────────────────────────┐
                 │           Bitstream          │
                 └─────┬──────────────────┬─────┘
                       │                  │
                       ▼                  ▼
                 ┌───────────┐      ┌───────────┐
                 │   CELT    │      │   SILK    │
                 │  decoder  │      │  decoder  │
                 └─────┬─────┘      └─────┬─────┘
                       │                  │
                       │                  ▼
                       │            ┌───────────┐
                       │            │   LACE    │
                       │            └─────┬─────┘
                       │                  │
                       │                  ▼
                       │            ┌───────────┐
                       │            │ Resampler │
                       │            └─────┬─────┘
                       │                  │
                       ▼                  ▼
                 ┌──────────────────────────────┐
                 │        Mode Handling         │
                 └──────────────┬───────────────┘
                                │
                                ▼
                         decoded  signal
      ]]>
        </artwork>
      </figure>

      <t>
        LACE has meanwhile been superseded by the non-linear adaptive coding enhancer (NoLACE) <xref target="nolace-paper"/> which shares all basic properties of
        LACE outlined above but provides higher quality. This stresses the advantage of specifying requirements for an enhancement algorithm over specifying the algorithm itself.
      </t>
    </section>

    <section anchor="enhancement-definition">
    <name>Definition of a SILK Enhancement Algorithm</name>
      <t>
      A SILK enhancement algorithm denotes any algorithm that replaces the native SILK decoded output with an enhanced output signal,
      an example of which is depicted in <xref target="opus-with-lace"/>. Here, the native SILK decoded output refers to the signal that an
      unmodified SILK decoder would produce. A SILK enhancement algorithm is not restricted to operating on the native SILK decoded output
      alone; it may equally draw on internal quantities of the SILK decoder (such as gains, LSFs, pitch or excitation), modify such internal
      values, or replace parts of the SILK decoder altogether. The requirements in this document therefore constrain only the relation between the
      native SILK decoded output and the output of the enhancement algorithm.
      If the encoded content is wideband and the decoder sampling rate allows for a higher bandwidth than wideband, a SILK enhancement algorithm
      may also produce an output signal of higher bandwidth than the native SILK decoded output, replacing the resampler in <xref target="opus-with-lace"/>,
      or it may modify the combination of a SILK-decoded wideband signal with a CELT decoded highband signal in hybrid mode. Such bandwidth
      extension is only permitted for wideband speech; narrowband and mediumband content MUST NOT be extended. In any case, a SILK enhancement
      algorithm MUST NOT modify the output of pure CELT frames. A SILK enhancement algorithm whose output has a higher bandwidth than the native SILK decoded output
      will be referred to as extending, whereas a SILK enhancement algorithm whose output preserves the bandwidth of the native SILK decoded
      output will be referred to as non-extending. Furthermore, an Opus decoder including a SILK enhancement algorithm will be referred to as
      enhanced decoder. Note however, that simply resampling the signal to a higher sampling rate is neither considered enhancing nor extending.
      </t>
    </section>

    <section>
    <name>Qualification of a SILK Enhancement Algorithm</name>

      <section>
      <name>General Requirements</name>
        <section anchor="enhancement-subjective">
          <name>Subjective Evaluation</name>
          <t>
            Objective metrics for quality evaluation have often proved unreliable especially for evaluating completely new algorithms for processing speech
            or audio signals. Therefore, any SILK enhancement algorithm SHOULD undergo subjective evaluation before integration into the Opus decoder. For genuinely
            new algorithms, it is RECOMMENDED to perform either an absolute category rating (ACR) or degradation category rating (DCR)
            listening test according to <xref target="p.800"/> or <xref target="p.808"/>, where the test conditions SHOULD cover a relevant
            range of bitrates. For modifications of previously tested algorithms, e.g. changing the size of a LACE model or adding small tunings for quality improvement or
            complexity reduction, at least an informal subjective evaluation SHOULD be carried out. Any enhancement algorithm SHOULD significantly improve quality for at least
            one encoder operating point while showing no significant degradation for other operating points.
          </t>
        </section>

        <section>
        <name>Delay and Phase Considerations</name>
          <t>
          SILK is approximately phase preserving and to avoid additional delay and maintain usability for applications relying on phase information, any SILK enhancement algorithm
          SHOULD also be approximately phase preserving.
          </t>
        </section>

        <section anchor="encoder-interop">
          <name>Encoder Requirements</name>
          <t>
            The Opus specification <xref target="RFC6716"/> provides much freedom for encoding an audio signal and the presence of a powerful enhancement
            algorithm can provide an incentive to use that freedom to produce bitstreams that, when decoded with a legacy Opus decoder, do not result in a reproduction
            of the input signal anymore. To prevent this, the following requirement is added for an Opus encoder that is designed to be used with an enhanced Opus
            decoder: if an Opus encoder produces a bitstream that can be decoded into a human-recognizable reproduction of the encoded signal with an enhanced Opus decoder, then
            that bitstream MUST also result in a human-recognizable reproduction of the encoded signal when decoded with a legacy Opus decoder.
          </t>
        </section>

        <section anchor="training-data">
          <name>Training Data</name>
          <t>
            To keep the objective tests meaningful, an enhancement algorithm MUST NOT be trained on any of the clips used in the tests of this document,
            and SHOULD NOT be trained on the datasets from which those clips are drawn (the Common Voice dataset for the wideband speech clips and
            additionally the VCTK corpus for the fullband speech clips).
          </t>
        </section>
      </section>

      <section>
      <name>Requirements Specific to Non-Extending SILK Enhancement Algorithms for Wideband Speech</name>

        <section anchor="nonext-enhancement-objective">
          <name>Lowband Test (Non-Extending)</name>
          <t>
            Every non-extending SILK enhancement algorithm for SILK decoded wideband speech signals MUST pass all objective tests put forth in this section.
            This collection of tests is designed to uncover major failure points of the tested algorithm that could be due to improper design or training data,
            or due to improper integration into the Opus decoder. It is not designed to (and cannot) assess the quality of a particular enhancement algorithm.
          </t>

          <t>
            The tests are based on comparing a degradation score for audio samples decoded from a list of bitstreams contained in
            <eref target="https://media.xiph.org/opus/ietf/osce_testvectors_v1.zip"/> (FIXME: find final location) to a reference degradation score computed from audio
            decoded with a reference decoder. Any conforming Opus decoder <xref target="RFC6716"/> is admissible as a reference decoder. This latitude is
            deliberate and necessary: decoding a bitstream may involve behavior that is not fully specified by <xref target="RFC6716"/>, such as packet loss
            concealment, so different conforming decoders may produce slightly different reference scores. Each test corresponds to an encoder operating point
            and the test names follow the scheme
          </t>
          <t>osce_test_BITRATE_BITRATEMODE_FRAMESIZEms_BANDWIDTH_cCOMPLEXITY_MODE</t>
          <t>where</t>
          <ol type="(%d)">
            <li>BITRATE is either a number specifying the encoder bitrate in
            bits per second or the string "SWITCHING" indicating the bitrate has been switched during encoding,</li>
            <li>BITRATEMODE is either vbr or cbr indicating variable bitrate or constant bitrate encoding,</li>
            <li>FRAMESIZE is either 10 or 20,</li>
            <li>BANDWIDTH specifies the maximal bandwidth and is always WB for this test (note however that the actual bandwidth can be lower),</li>
            <li>COMPLEXITY is a number from 0 to 10 and specifies the encoder complexity,</li>
            <li>MODE refers to the coding mode and is either "native" or "celtswitching". In "native" mode, the encoder decision whether to use SILK or CELT is
            based on signal classification whereas in "celtswitching" mode the encoder has been forced to switch between SILK and CELT at a fixed rate.</li>
          </ol>
          <t>
            The testvectors are further divided into groups, where each group contains either speech samples from the same language or dialect, or music
            content. Each group GROUP is tested separately and the test is passed if it is passed for every group. The bitstreams in TESTNAME/GROUP follow
            the naming pattern CLIPNAME_TESTNAME which associates each bitstream uniquely with a reference signal reference_clips/CLIPNAME.s16. For every
            CLIPNAME in GROUP let REFMOC(CLIPNAME) denote the reference degradation score, obtained by decoding the bitstream with a reference decoder and
            scoring the result against reference_clips/CLIPNAME.s16 in the same way as the test signal below. Furthermore, let
            CLIPNAME_test.s16 denote the signal decoded with the enhanced decoder under test
            at a sampling frequency of 16 kHz after delay compensation. The degradation for the test signal CLIPNAME_test.s16 is calculated using the osce_compare tool
            <eref target="https://gitlab.xiph.org/xiph/opus/-/blob/osce-testing/src/osce_compare.c"/> with the reference signal path as first argument and the test signal path as second argument. The resulting degradation score will be referred to as TESTMOC(CLIPNAME).
          </t>
          <t>
            All degradation and distortion scores in this document MUST be computed on time-aligned signals: before scoring, the decoder delay MUST
            be compensated so that the decoded signal is sample-aligned with the reference signal reference_clips/CLIPNAME.s16. The osce_compare tool
            performs this compensation by discarding a fixed number of leading samples from the decoded (second-argument) signal, given through its
            -delay option. The decoder delay depends on the Opus implementation and its configuration, and, for an enhanced decoder, on the SILK
            enhancement algorithm; it may therefore differ between the reference decoder and the decoder under test as well as between sampling rates.
            The applicable delay MUST be determined for each decoder and sampling rate used, since residual misalignment degrades the scores and can
            invalidate the test result.
          </t>
          <t>From the reference degradation score REFMOC(CLIPNAME) and the test degradation score TESTMOC(CLIPNAME) a difference score is calculated according to</t>

            <artwork type="ascii-art" name="moc_diff.txt">
          <![CDATA[
                    REFMOC(CLIPNAME) - TESTMOC(CLIPNAME)
      D(CLIPNAME) = ------------------------------------
                                               0.5
                        0.1 + REFMOC(CLIPNAME)
      ]]>
            </artwork>

          <t>To pass the test for group GROUP, the following two criteria MUST be met:</t>
            <ol type="(%d)">
            <li>D(CLIPNAME) is larger than A for every CLIPNAME in GROUP,</li>
            <li>The average of D(CLIPNAME) over GROUP is larger than B.</li>
          </ol>
          <t>
          The thresholds are A = -0.5 and B = -0.052. A test is passed if it is passed for all groups in that test.
          </t>
        </section>
      </section>
        <section anchor="ext-enhancement-requirements">
        <name>Requirements Specific to Extending SILK Enhancement Algorithms for Wideband Speech</name>
        <t>
          An extending SILK enhancement algorithm extends the bandwidth of the decoded speech beyond the 8 kHz upper limit of wideband. The objective
          tests in this section therefore evaluate the decoder output at a sampling frequency of 48 kHz. The blind bandwidth extension network (BBWENet)
          <xref target="bbwenet-paper"/> is an example of an extending SILK enhancement algorithm.
        </t>
        <t>
          An extending algorithm is subject to two objective tests:
        </t>
        <ul>
          <li>a lowband test, which ensures that the quality of the 0-8 kHz lowband is not degraded relative to an unenhanced decoder; and</li>
          <li>a highband test, which checks that the extended signal has a closer per-band spectral distortion to the fullband reference than the
          unenhanced wideband-decoded signal has to that reference.</li>
        </ul>
        <t>
          The lowband test is the objective test of <xref target="nonext-enhancement-objective"/> applied to the 48 kHz output; a non-extending
          algorithm is subject only to that test. An extending algorithm MUST pass the lowband test.
        </t>
        <t>
          Bandwidth extension is blind: the highband content above 8 kHz is not present in the native SILK decoded output, and many plausible highband extensions
          exist, some of which may violate the highband criterion without being perceptually wrong. The highband test is therefore relaxed in two ways:
          the criterion is only required to hold for most of the test clips rather than every clip, and an extending algorithm SHOULD, rather than MUST,
          pass the highband test. Subjective evaluation as described in <xref target="enhancement-subjective"/> remains the primary means of assessing the
          highband of an extending algorithm.
        </t>

          <section anchor="ext-lowband">
            <name>Lowband Test (Extending)</name>
            <t>
              The lowband test reuses the testvectors, naming scheme, grouping, difference score D(CLIPNAME) and group pass criteria of
              <xref target="nonext-enhancement-objective"/>, with the following modifications:
            </t>
            <ol type="(%d)">
              <li>the reference decode (yielding REFMOC(CLIPNAME)) is produced by decoding the bitstream with the reference decoder at a sampling
              frequency of 48 kHz instead of at 16 kHz;</li>
              <li>the test signal (yielding TESTMOC(CLIPNAME)) is produced by decoding the bitstream with the enhanced decoder under test, i.e. the
              decoder applying the extending enhancement algorithm, at a sampling frequency of 48 kHz;</li>
              <li>both degradation scores are computed on the 0-8 kHz lowband only.</li>
            </ol>
            <t>
              Both degradation scores are obtained with the osce_compare tool
              <eref target="https://gitlab.xiph.org/xiph/opus/-/blob/osce-testing/src/osce_compare.c"/>, invoked with the 16 kHz reference signal
              reference_clips/CLIPNAME.s16 as first argument and the respective 48 kHz decoded signal as second argument, with the reference and test
              sampling rates set accordingly (-fs_ref 16000 -fs_test 48000) and with the applicable delay compensation (-delay). Given these rates,
              osce_compare reduces the 48 kHz decoded signal to the 0-8 kHz lowband before scoring it against the reference signal. Decoding both the
              reference decode and the test signal at 48 kHz ensures that the 16 kHz to 48 kHz resampling effects on the lowband cancel between the two
              in the difference score D(CLIPNAME).
            </t>
            <t>
              In all other respects the lowband test is identical to <xref target="nonext-enhancement-objective"/>. An extending algorithm MUST pass
              the lowband test, i.e. it MUST NOT degrade the wideband content.
            </t>
          </section>

          <section anchor="ext-highband">
            <name>Highband Test</name>
            <t>
              The highband test verifies that the algorithm adds meaningful signal energy above 8 kHz and does not degrade the highband.
              It is based on a separate set of fullband (48 kHz) speech reference clips contained in
              <eref target="https://media.xiph.org/opus/ietf/osce_testvectors_v1.zip"/> (FIXME: find final location) under the highband subdirectory.
              Each reference clip is encoded as wideband and the resulting bitstream is decoded with the extending decoder under test at a sampling
              frequency of 48 kHz. Each test corresponds to an encoder operating point and the test names follow the scheme
            </t>
            <t>osce_hbtest_BITRATE_BITRATEMODE_FRAMESIZEms_BANDWIDTH_cCOMPLEXITY_MODE</t>
            <t>with the fields defined as in <xref target="nonext-enhancement-objective"/>. The bitstreams in TESTNAME follow the naming pattern
              CLIPNAME_TESTNAME which associates each bitstream uniquely with a fullband reference signal highband/reference_clips/CLIPNAME.s16.</t>
            <t>
              For a decoded test signal and its fullband reference, a per-band spectral distortion DIST(reference, test) is computed for each of the four
              highbands spanning 8.0-9.6, 9.6-12.0, 12.0-15.6 and 15.6-20.0 kHz, together with the distortion DIST(reference, anchor) of a lowpass
              anchor. The lowpass anchor is the "no extension" signal and is derived from the fullband reference by zeroing out all STFT bins at and above 8 kHz. The exact computation is defined by the osce_compare tool
              <eref target="https://gitlab.xiph.org/xiph/opus/-/blob/osce-testing/src/osce_compare.c"/>, invoked with its -highband option. A clip passes
              when, in every one of the four highbands BAND, the extended signal is at least as close to the reference as the anchor:
            </t>

            <artwork type="ascii-art" name="highband_criterion.txt">
          <![CDATA[
      DIST(reference, test)(BAND) <= DIST(reference, anchor)(BAND)
      ]]>
            </artwork>
            <t>
              As with the lowband tests (see <xref target="nonext-enhancement-objective"/>), the decoded test signal MUST be delay-compensated so
              that it is sample-aligned with the fullband reference highband/reference_clips/CLIPNAME.s16 before the distortions are computed. The
              osce_compare tool applies this compensation through its -delay option at the 48 kHz test sampling rate; the applicable decoder delay
              depends on the Opus implementation and the enhancement algorithm and may differ between decoders.
            </t>
            <t>
              For each test (encoder operating point) the pass rate is the fraction of clips that pass, and the test is passed if more than 90% of the
              clips pass. The highband test is only meaningful when the lowband quality is sufficiently high, which requires a sufficiently high
              bitrate for the given frame size. Therefore, in keeping with the blind nature of the problem described above, an extending algorithm
              SHOULD pass all tests for bitrates of 9 kb/s and above when the frame size is 20 ms, and it SHOULD pass all tests for bitrates of
              12 kb/s and above when the frame size is 10 ms. Lower bitrates are included in the testvectors, but the test result on those operating
              points is informative only.
            </t>
          </section>
        </section>
    </section>

    <section anchor="IANA">
    <!-- All drafts are required to have an IANA considerations section. See RFC 8126 for a guide.-->
      <name>IANA Considerations</name>
      <t>
        This document updates the media type registration for the audio/opus media subtype defined in <xref target="RFC7587"/> by adding the two optional parameters below to its list of optional parameters. Both are informative
        parameters by which a receiver signals its decoder's intent to enhance (and, for extendedbandwidth, to extend) the received signal, so that the
        far-end encoder MAY select an operating point (for example, a lower bitrate or a wideband-only mode) that the enhancement can compensate for. The
        decoder is not bound by this signaling and may at runtime reduce or disable enhancement, for example when CPU load is high. The parameters do not
        change which bitstreams a decoder accepts, and by the requirement of <xref target="encoder-interop"/> any stream that relies on them MUST remain
        intelligible when decoded by a legacy Opus decoder.
      </t>
      <dl>
        <dt>speechenhancement:</dt>
        <dd>
          an integer from 0 to 40 indicating the effectiveness of the receiver's speech enhancement (see <xref target="enhancement-definition"/>). It is
          calibrated at 6 kb/s: if the enhanced decoder at 6 kb/s achieves the quality of an unenhanced decoder at X kb/s, the value is
          round(10 * (X - 6) / 6). A value of 0 (the default when the parameter is absent) indicates no speech enhancement; 10 indicates that
          enhancement at 6 kb/s reaches the quality of an unenhanced decoder at 12 kb/s, and 40 (the maximum) corresponds to the quality of an
          unenhanced decoder at 30 kb/s or better.
        </dd>
        <dt>extendedbandwidth:</dt>
        <dd>
          an integer giving the upper limit, in kHz, of the audio bandwidth the receiver's decoder produces after blind bandwidth extension. It applies only
          to content coded in the speech coding mode; content coded in the general (music) coding mode is not affected by this parameter and is reproduced at
          the bandwidth selected by the encoder. Typical values follow the Opus bandwidth tiers: 8 (wideband, i.e. no extension and the default when the
          parameter is absent), 12 (super-wideband) and 20 (fullband). For example, an extending algorithm such as BBWENet <xref target="bbwenet-paper"/> that
          extends wideband speech to fullband would signal extendedbandwidth=20.
        </dd>
      </dl>
      <t>
        An offerer or answerer MAY include either parameter; a value that is absent, zero (speechenhancement) or 8 (extendedbandwidth) is equivalent to the
        corresponding baseline capability. An implementation that does not recognize these parameters ignores them, as required by <xref target="RFC7587"/>.
      </t>
    </section>

    <section anchor="Security">
      <!-- All drafts are required to have a security considerations section. See RFC 3552 for a guide. -->
      <name>Security Considerations</name>
      <t>
        A SILK enhancement algorithm runs after the SILK decoder and post-processes (or replaces) its output. It introduces no new bitstream
        elements and does not change which packets a decoder accepts or rejects, so it does not expand the packet-parser attack surface and the
        security considerations of <xref target="RFC6716"/> apply unchanged.
      </t>
      <t>
        The following additional considerations apply to the enhancement itself:
      </t>
      <ul>
        <li>Memory safety on attacker-influenced input: the enhancement operates on the decoder output and on internal parameters (such as gains,
        LSFs, pitch and excitation) that are ultimately derived from untrusted bitstreams. An implementation MUST handle degenerate or extreme
        values, including out-of-range or non-finite intermediates, without out-of-bounds access, and MUST run in bounded memory.</li>
        <li>Bounded, data-independent complexity: the example enhancement models (LACE, NoLACE and BBWENet) are of fixed size, so their per-frame
        cost is essentially constant and independent of the packet content. Future enhancement models SHOULD likewise avoid any data-dependent
        unbounded work, so that crafted streams cannot force excessive CPU usage; because the enhancement is gated by decoder complexity, a decoder
        can also bound or disable it.</li>
        <li>Model integrity: the model weights are part of the implementation and are not signalled in the bitstream, so they are not
        attacker-controlled through Opus packets. When the weights are loaded from an external artifact, their integrity should be ensured (for
        example through a verified hash). A corrupted model can only degrade audio quality, not compromise memory safety, provided inference is
        bounds-checked.</li>
      </ul>
    </section>

    <!-- NOTE: The Acknowledgements and Contributors sections are at the end of this template -->
  </middle>

  <back>
    <references>
      <name>References</name>
      <references>
        <name>Normative References</name>
        <xi:include href="https://www.rfc-editor.org/refs/bibxml/reference.RFC.2119.xml"/>
        <xi:include href="https://www.rfc-editor.org/refs/bibxml/reference.RFC.8174.xml"/>
        <xi:include href="https://www.rfc-editor.org/refs/bibxml/reference.RFC.6716.xml"/>
        <xi:include href="https://www.rfc-editor.org/refs/bibxml/reference.RFC.7587.xml"/>
        <reference anchor="p.800" target="https://www.itu.int/rec/T-REC-P.800-199608-I">
          <front>
            <title>P.800 : Methods for subjective determination of transmission quality</title>
            <author><organization abbrev="ITU-T">International Telecommunication Union - Telecommunication Standardization Sector</organization></author>
            <date month="August" year="1996"/>
          </front>
        </reference>
        <reference anchor="p.808" target="https://www.itu.int/rec/T-REC-P.808-202106-I/en">
          <front>
            <title>P.808 : Subjective evaluation of speech quality with a crowdsourcing approach</title>
            <author><organization abbrev="ITU-T">International Telecommunication Union - Telecommunication Standardization Sector</organization></author>
            <date month="June" year="2021"/>
          </front>
        </reference>
      </references>

      <references>
        <name>Informative References</name>
        <reference anchor="lace-paper" target="https://doi.org/10.1109/WASPAA58266.2023.10248150">
          <front>
            <title>LACE: A light-weight, causal Model for enhancing coded Speech through Adaptive Convolutions</title>
            <author initials="J." surname="Buethe"/>
            <author initials="J.-M." surname="Valin"/>
            <author initials="A." surname="Mustafa"/>
            <date year="2023"/>
          </front>
          <refcontent>2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, pp. 1-5</refcontent>
          <seriesInfo name="DOI" value="10.1109/WASPAA58266.2023.10248150"/>
        </reference>
        <reference anchor="nolace-paper" target="https://doi.org/10.1109/ICASSP48485.2024.10448332">
          <front>
            <title>NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping</title>
            <author initials="J." surname="Buethe"/>
            <author initials="A." surname="Mustafa"/>
            <author initials="J.-M." surname="Valin"/>
            <author initials="K." surname="Helwani"/>
            <author initials="M." surname="Goodwin"/>
            <date year="2024"/>
          </front>
          <refcontent>ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, pp. 476-480</refcontent>
          <seriesInfo name="DOI" value="10.1109/ICASSP48485.2024.10448332"/>
        </reference>
        <reference anchor="bbwenet-paper" target="https://doi.org/10.1109/WASPAA66052.2025.11230955">
          <front>
            <title>A lightweight and robust method for blind wideband-to-fullband extension of speech</title>
            <author initials="J." surname="Buethe"/>
            <author initials="J.-M." surname="Valin"/>
            <date year="2025"/>
          </front>
          <refcontent>2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), Tahoe City, CA, USA, pp. 1-5</refcontent>
          <seriesInfo name="DOI" value="10.1109/WASPAA66052.2025.11230955"/>
        </reference>
      </references>
    </references>


 </back>
</rfc>
