Skip to content

olo_fix_fir_dec_semi_chtdm

Back to Entity List

Status Information

VHDL Source: olo_fix_fir_dec_semi_chtdm.vhd
Bit-true Model: olo_fix_fir_dec.py

Description

This entity implements a decimating FIR filter for one or more TDM (time-division-multiplexed) channels. All channels share the same coefficient set and are processed one after the other. The filter taps are computed semi-parallel: Multipliers_g multiply-add operations are chained together in the classic MACC (multiply-accumulate) chain and ceil(Taps_g / Multipliers_g) clock cycles are used to compute one output sample for one channel.

A single channel (Channels_g = 1) is supported as well; the input is then a plain sample stream.

The number of multipliers therefore trades resources against throughput:

  • Multipliers_g = 1 behaves like a fully serial FIR (one tap per clock cycle).
  • Multipliers_g = Taps_g behaves like a fully parallel FIR (all taps in one clock cycle).
  • Any value in between provides a semi-parallel implementation.

Example: A 4 channel, 16 taps FIR filter with Multipliers_g = 4 requires 4 x 4 = 16 clock cycles to produce one output sample set (one sample for each channel).

Note that the filter can also be used without decimation (Ratio_g = 1).

For details about the fixed-point number format used in Open Logic, refer to the fixed point principles.

Coefficients can be fixed (ROM) or runtime configurable (RAM) with optional readback.

Input Bandwidth Limitation

This entity does not generate backpressure. The semi-parallel MACC chain requires ceil(Taps_g / Multipliers_g) x Channels_g clock cycles to compute one output sample set (for all channels). This calculation is repeated every Ratio_g input sample sets.

fin≤fclk·Ratio_g·Multipliers_gTapsg·Channelsg

where f_in is the rate of complete TDM frames (one frame = Channels_g samples). If the input arrives faster than this limit, the filter stops working correctly. In simulation an error is reported if the processing power is insufficient.

If data is arriving faster than the filter can process, an error message is reported in simulation.

Use olo_base_rate_limit externally to enforce the rate limit.

Unless FullInpRateSupport_g = true, at least one idle cycle (In_Valid = '0') is required between two consecutive input samples. See Full Input Rate Support.

Latency

This block changes the sample rate. Because not every input sample produces an output sample, the latency is not fixed and is therefore not documented in detail.

Generics

General Generics

Name Type Default Description
InFmt_g string - Input format
String representation of an en_cl_fix FixFormat_t
OutFmt_g string - Output format
String representation of an en_cl_fix FixFormat_t
CoefFmt_g string - Coefficient format
String representation of an en_cl_fix FixFormat_t
Channels_g positive 1 Number of TDM channels (>= 1; single- or multi-channel)
Ratio_g positive 1 Decimation ratio (one output per Ratio_g input sample sets)
Taps_g positive - Number of filter taps (must be >= 2)
Multipliers_g positive - Number of multipliers (MACC-chain lanes) computed in parallel
FullInpRateSupport_g boolean false true - input samples may be applied on every clock cycle (uses an additional delay line).
false - at least one idle cycle is required between input samples.
GuardBits_g natural 1 Number of integer guard bits in the accumulator above OutFmt_g
Round_g string "Trunc_s" Rounding mode
String representation of an en_cl_fix FixRound_t
Saturate_g string "Warn_s" Saturation mode
String representation of an en_cl_fix FixSaturate_t
MultRegs_g positive 1 Number of pipeline registers in each multiplier

By nature a semi-parallel FIR filter makes sense only for input data-rates below one sample per clock cycle (otherwise a fully parallel FIR is more efficient). However, the generic FullInpRateSupport_g=True can be important for bursty inputdata streams. But often it is more resource efficient to use FullInpRateSupport_g=False and throttle the input data rate by using a olo_base_rate_limit entity in front of the filter.

Coefficient and Data Storage

Name Type Default Description
CoefInit_g string "0.0" Comma-separated initial coefficient values (real numbers, quantized to CoefFmt_g)
Example: "0.3, 0.55, 0.2"
see olo_fix_coef_storage
CoefStorageType_g string "ROM" Coefficient storage type: "ROM" (fixed) or "RAM" (runtime-updateable)
see olo_fix_coef_storage
CoefRamReadback_g boolean false Enable coefficient readback via Coef_Rd_... ports (RAM mode only)
see olo_fix_coef_storage
CoefRamBehavior_g string "RBW" Coefficient RAM behavior: "RBW" = read-before-write, "WBR" = write-before-read
see olo_fix_coef_storage
CoefMemStyle_g string "auto" Synthesis attribute for coefficient memory style (e.g. "block", "distributed")
see olo_fix_coef_storage
DataRamBehavior_g string "RBW" Data RAM behavior: "RBW" = read-before-write, "WBR" = write-before-read
see olo_base_ram_tdp
DataMemStyle_g string "auto" Synthesis attribute for data RAM style (e.g. "block", "distributed")
see olo_base_ram_tdp

Interfaces

Control

Name In/Out Length Default Description
Clk in 1 - Clock
Rst in 1 - Reset (synchronous, active high)

Coefficient Configuration

Name In/Out Length Default Description
Coef_Addr in log2ceil(Taps_g) 0 Coefficient address for read/write
Coef_WrEna in 1 '0' Coefficient write enable (RAM mode only)
Coef_WrData in width(CoefFmt_g) 0 Coefficient write data (RAM mode only)
Coef_RdEna in 1 '0' Coefficient read enable (RAM readback mode only)
Coef_RdData out width(CoefFmt_g) N/A Coefficient read data (0 in ROM mode)
Coef_RdValid out 1 N/A Coefficient read valid (0 in ROM mode)

All Coef_* ports have safe defaults and can be left unconnected in ROM mode or when coefficient updates are not needed.

Delay-Line Flushing

Name In/Out Length Default Description
Flush_Ena in 1 '0' A pulse on this port starts a flush that zeros all data delay lines.
Flush_Done out 1 N/A A pulse on this port indicates that a flush started by Flush_Ena finished.

See Startup and Flushing.

Input Data

Name In/Out Length Default Description
In_Valid in 1 - Input valid
In_Data in width(InFmt_g) - Input data (TDM: channels interleaved, ch0 first)
In_Last in 1 '0' TDM frame boundary (optional)
see TDM Conventions

The In_Last signal is optional and has no functional effect. In simulation it is only used to check that it is asserted at the correct TDM position (last channel); an error is reported if In_Last is asserted on a sample of any other channel. See Last Handling.

Output Data

Name In/Out Length Default Description
Out_Valid out 1 N/A Output valid
Out_Data out width(OutFmt_g) N/A Output data (TDM: channels interleaved, ch0 first)
Out_Last out 1 N/A TDM frame boundary, asserted on the last channel
see TDM Conventions

Details

Example Instantiation

The example below shows a simple instantiation: fixed coefficients stored in ROM, four taps computed with two multipliers and a decimation ratio of two. All coefficient configuration ports, the flushing interface and In_Last are omitted.

i_fir : entity olo.olo_fix_fir_dec_semi_chtdm
    generic map (
        -- Formats
        InFmt_g       => "(1,0,15)",
        OutFmt_g      => "(1,0,15)",
        CoefFmt_g     => "(1,0,17)",
        -- Filter parameters
        Channels_g    => 4,
        Ratio_g       => 2,
        Taps_g        => 4,
        Multipliers_g => 2,
        -- Fixed coefficients stored in ROM
        CoefInit_g    => "0.1, 0.4, 0.4, 0.1"
    )
    port map (
        Clk       => Clk,
        Rst       => Rst,
        In_Valid  => In_Valid,
        In_Data   => In_Data,
        Out_Valid => Out_Valid,
        Out_Data  => Out_Data
    );

Architecture

The datapath consists of Multipliers_g parallel MACC stages built from olo_fix_madd. The stages are chained (each stage adds its product to the running sum coming from the previous stage) so the synthesizer can map them onto a DSP cascade. In every calculation cycle the chain produces the sum of Multipliers_g tap products; these partial sums are accumulated over ceil(Taps_g / Multipliers_g) cycles to form one output sample.

This architecture is depicted by below example of a 2-stage architecture (Multipiers_g = 2):

architecture

Each stage owns:

  • A data delay-line RAM (olo_base_ram_tdp). The RAMs are chained so that each stage sees the input delayed by a further block of taps.
  • A coefficient storage (olo_fix_coef_storage) holding only that lane's block of ceil(Taps_g / Multipliers_g) coefficients (ROM or RAM depending on CoefStorageType_g). The coefficient memory is therefore split across the stages rather than replicated, so the total coefficient memory does not grow with Multipliers_g. Coefficient writes and readbacks addressed through the Coef_... ports are routed to (and muxed back from) the stage that owns the addressed tap.

Below figure depicts a single stage of the MACC chain.

stage

The result of the accumulation is rounded and saturated to OutFmt_g using olo_fix_resize.

The memory styles of the coefficient storage and the data RAM can be selected independently through CoefMemStyle_g and DataMemStyle_g. To minimize coefficient memory, choose a coefficient storage type that fits the use case: ROM for fixed coefficients, or RAM without readback (CoefRamReadback_g = false) when runtime updates are needed but readback is not.

Full Input Rate Support

When FullInpRateSupport_g = false (default), the chained data memory needs one idle cycle between two input samples, hence In_Valid must not be asserted on two consecutive clock cycles.

When FullInpRateSupport_g = true, an additional delay line (olo_base_delay) per stage provides the chained delay, so In_Valid may be asserted on every clock cycle. Note that this only relaxes the back-to-back input restriction; the overall processing power limit (see Input Bandwidth Limitation) still applies - and the price for the architecture is additional memory for the extra delay lines. Whenever possible it is to be preferred to avoid In_Valid being asserted on consecutive clock cycles and to use FullInpRateSupport_g = false.

The stage architecture with FullInpRateSupport_g = true is depicted below:

full-input-rate-support

Startup and Flushing

The delay lines are stored in RAM and are not cleared by reset. After power-up the RAMs are zero initialized, hence the first outputs are bit-true without any special action. If the filter is reset during operation, the RAMs still contain old data. In this case a flush must be triggered (pulse Flush_Ena and wait for Flush_Done) to zero the delay lines before feeding new data. This guarantees bit-true agreement with the Python model, which initializes its delay line to zero.

Coefficient Format

The accumulator operates at full multiply precision:

  • MultFmt = (max(In.S, Coef.S), In.I + Coef.I, In.F + Coef.F)
  • AccuFmt = (1, Out.I + GuardBits_g, In.F + Coef.F) (GuardBits_g guard bits above output)

Choosing OutFmt.I or GuardBits_g too small risks accumulator overflow. Ensure max_sum_of_products <= 2^(OutFmt.I + GuardBits_g) - 1 LSB.

Accumulator Guard Bits

The accumulator carries GuardBits_g integer guard bits above OutFmt_g (AccuFmt.I = OutFmt.I + GuardBits_g). These bits allow the sum of products to grow beyond the output range during the accumulation without overflowing. With the default of one guard bit, intermediate results of up to twice the OutFmt_g maximum are supported. The user is responsible for choosing GuardBits_g, the coefficients and the formats such that the accumulator does not overflow; otherwise the number of guard bits or the output format must be increased.

Last Handling

On the input, In_Last is not required for operation. It is only used in simulation to detect incorrect TDM framing: an error is reported if In_Last is asserted on a sample that does not belong to the last channel (Channels_g-1). It has no functional effect on the computation.

On the output, Out_Last is generated by the entity itself and is always asserted together with Out_Valid on the last channel (Channels_g-1) of every output sample set.