CSIT/TestFailuresTracking

From fd.io
< CSIT
Revision as of 12:36, 12 July 2023 by Vrpolak (Talk | contribs)

Jump to: navigation, search

Contents

CSIT Test Failure Clasification

All known CSIT failures grouped and listed in the following order:

  • Always failing followed by sometimes failing.
  • Always failing tests:
    • Most common use cases followed by less common.
  • Sometimes failing tests:
    • Most frequently failing followed by less frequently failing.
      • High frequency 50%-100%
      • medium frequency 10%-50%
      • low frequency 0%-10%.
    • Within each sub-group: most common use cases followed by less common.

CSIT Test Fixing Priorities

Test fixing work priorities defined as follows:

  • (H)igh priority, most common use cases and most common test code.
  • (M)edium priority, specific HW and pervasive test code issue.
  • (L)ow priority, corner cases and external dependencies.

Current Failures

Deterministic Failures

In Trending

(M) 3n-snr: All hwasync wireguard tests failing when trying to verify device

  • last update: before 2023-01-31
  • work-to-fix: hard
  • rca: Missing QAT driver. Symptom: Failed to bind PCI device 0000:f4:00.0 to c4xxx on host 10.30.51.93
  • test: hwasync wireguard
  • frequency: always
  • testbed: 3n-snr
  • example: 3n-snr
  • ticket: CSIT-1883

(M) DPDK 23.03 testpmd startup fails on some testbeds

(M) 2n-spr: zero traffic on cx7 rdma

(M) 3n-icx, 3n-snr: first few swasync scheduler tests timing out in runtime stat

(L) 3n-icx: negative ipackets on TB38 AVF 4c l2patch

(L) 2n-tx2: af_xdp mrr failures

Not In Trending

(M) all testbeds: some 9000B tests

  • last update: 2023-02-09
  • work-to-fix: hard
  • rca: VPP code: 34839: dpdk: cleanup MTU handling. CSIT needs to rework how it sets MTU / max frame rate (CSIT-1797). Some tests will continue failing due to missing support on VPP side, we will open specific Jira tickets for those.
  • test: see sub-items
  • frequency: always
  • testbed: all
  • examples: see sub-items
  • ticket: CSIT-1809
  • gerrit: https://gerrit.fd.io/r/c/csit/+/37824
(M) tests with 9000B payload frames not forwarded over vhost interfaces
  • last update: 2023-02-09
  • work-to-fix: hard
  • test: 9000B + vhostuser
  • testbed: 2n-skx, 3n-skx, 2n-clx
  • examples: 3n-skx vhostuser
  • ticket: CSIT-1809
(M) tests with 9000B payload frames not forwarded over memif interfaces
  • last update: 2023-02-09
  • work-to-fix: hard
  • test: 9000B + memif
  • testbed: 2n-skx, 3n-skx, 2n-clx
  • examples: 2n-skx Memif
  • ticket: CSIT-1808
(M) 9000B payload frames not forwarded over tunnels due to violating supported Max Frame Size (VxLAN, LISP, SRv6)
  • last update: 2023-02-09
  • work-to-fix: medium
  • test: 9000B + (IP4 tunnels VXLAN, IP4 tunnels LISP, Srv6, IpSec)
  • testbed: 2n-icx, 3n-icx
  • examples: 2n-icx VXLAN, 3n-icx
  • ticket: CSIT-1801
(M) 9000B all AVF tests are failing to forward traffic
  • last update: 2023-02-09
  • work-to-fix: hard
  • test: 9000B + AVF
  • testbed: 3n-icx
  • examples: 3n-icx ip4base
  • ticket: CSIT-1885
(L) l3fwd error in 200Ge2P1Cx7Veat-Mlx5 test with 9000B
  • last update: 2023-06-28
  • work-to-fix: medium
  • test: 9000B + Cx7 with DPDK DUT
  • testbed: 2n-icx
  • examples: [1]
  • ticket: CSIT-1924

(M) 2n-clx, 2n-icx: all Geneve tests with 1024 tunnels fail

  • last update: before 2023-01-31
  • work-to-fix: hard
  • rca: VPP crash, Failed to add IP neighbor on interface geneve_tunnel258
  • test: avf-ethip4--ethip4udpgeneve-1024tun-ip4base 64B 1518B IMIX 1c 2c 4c
  • frequency: always
  • testbed: 2n-skx, 2n-clx, 2n-icx
  • example: 2n-icx
  • ticket: CSIT-1800

(L) 2n-clx, 2n-icx: nat44ed cps 16M sessions scale fail

  • last update: before 2023-01-31
  • work-to-fix: hard
  • rca: VPP crash, Failed to set NAT44 address range on host 10.30.51.44 (connections-per-second tests only)
  • test: 64B-avf-ethip4tcp-nat44ed-h262144-p63-s16515072-cps-ndrpdr 1c 2c 4c, 64B-avf-ethip4udp-nat44ed-h262144-p63-s16515072-cps-ndrpdr 1c 2c 4c
  • frequency: always
  • testbeds: 2n-skx, 2n-clx, 2n-icx
  • example: 2n-icx, 2n-clx
  • ticket: CSIT-1799

(L) 2n-clx, 2n-icx: nat44det imix 1M sessions fails to create sessions

Occasional Failures

In Trending

(H) 2n-icx: NFV density VPP does not start in container

  • last update: before 2023-01-31
  • work-to-fix: hard
  • rca:
  • test: all subsequent
  • frequency: medium
  • testbed: 2n-icx
  • example: 2n-icx mrr, 2n-icx ndrpdr
  • ticket: CSIT-1881
  • note: Once VPP breaks, all subsequent tests fail. Even all subsequent builds will be failing until Peter makes TB working again. Although it's failing with medium frequency when it happens it breaks all subsequent builds on the TB therefore [H] priority.

(M) 2n-clx: e810 mlrsearch tests packets forwarding in one direction

(M) 3n-icx, 3n-snr: wireguard 100 and 1000 tunnels mlrsearch tests failing with 2c and 4c

  • last update: before 2023-01-31
  • work-to-fix: easy
  • rca:
  • test: wireguard 100 tunnels and more
  • frequency: high
  • testbed: 3n-icx, 3n-snr
  • examples: 3n-icx
  • ticket: CSIT-1886

Rare Failures

In Trending

(M) 3n-icx, 3n-snr: 1518B IPsec packets not passing

(M) all testbeds: mlrsearch fails to find NDR rate

  • last update: before 2023-06-22
  • work-to-fix: hard
  • rca: On 3n-tsh, the symptom is TRex reporting ierrors, only in one direction. Other testbeds may have a different symptom, but failures there are less frequent.
  • test: Crypto, Ip4, L2, Srv6, Vm Vhost (all packet sizes, all core configurations affected)
  • frequency: low
  • testbed: 3n-tsh, 3n-alt, 2n-clx
  • example: [2]
  • ticket: CSIT-1804

(M) all testbeds: AF_XDP mlrsearch fails to find NDR rate

  • last update: before 2023-01-31
  • work-to-fix: hard
  • rca:
  • test: af-xdp multicore tests
  • frequency: low
  • testbed: 2n-clx, 2n-skx, 2n-tx2, 2n-icx
  • example: 2n-skx, 2n-clx
  • ticket: CSIT-1802
  • note: This is mainly observed in iterative and coverage. It's very low frequency ~ 1 out of 100

(L) all testbeds: vpp create avf interface failure in multi-core configs

  • last update: 2023-02-06
  • work-to-fix: hard
  • rca: issue in Intel FVL driver
  • test: multicore AVF
  • frequency: low
  • testbed: all testbeds
  • example: 2n-clx, 3n-icx
  • ticket: CSIT-1782
  • note: A long standing issue without a final permanent fix.

(L) all testbeds: nat44det 4M and 16M scale 1 session not established

  • last update: 2023-02-14
  • work-to-fix: hard
  • rca: unknown
  • test: nat44det udp 4m and 16m (64k is ok, 1m can fail but rarely than bigger scales)
  • frequency: low
  • testbed: 2n-zn2, 2n-skx, 2n-icx, 2n-clx
  • example: 2n-zn2, 2n-clx
  • ticket: CSIT-1795

Past Failures

(H) AVF suite setup fails if previous suite was also AVF

(M) 2n-spr 200Ge2P1Cx7Veat: TRex sees port line rate as 100 Gbps

(M) 2n-icx: interface down in nginx tests

(M) hoststack: ip4udpscale1cl10s-ldpreload-iperf3 times out waiting for strace

(M) 3n-tsh: vpp in VM starting too slowly

  • last update: before 2023-02-22
  • work-to-fix: medium
  • rca: perhaps related to numa, investigation continues
  • test: 3n-tsh: sporadic VM vhost
  • frequency: high
  • testbed: 3n-tsh
  • example: 3n-tsh, 3n-tsh
  • ticket: CSIT-1877
  • note: Applied a workaround that simply waits longer. May lead to a rare failure, but for now considered fixed.

(M) 2n-clx, 2n-icx, 2n-zn2: DPDK testpmd 9000b tests on xxv710 nic are failing with no traffic

  • last update: 2023-06-28
  • work-to-fix: medium
  • rca: The DPDK app only attempts to set MTU once, but if interface is down (CSIT-1848) it fails. As a workaround, MTU could be set on Linux interface before starting the DPDK app.
  • test: DPDK testpmd 9000b
  • frequency: always
  • testbed: 2n-clx, 2n-icx, 2n-zn2
  • example: 2n-clx, 2n-icx
  • ticket: CSIT-1870
  • note: Fixed, probably by the CSIT MTU handling change.

(M) 3n-alt: testpmd no traffic forwarded

  • last update: 2023-06-29
  • work-to-fix: medium
  • rca: DUT-DUT link takes too long to come up on some testbeds. This happens *after* a test case with a DPDK app (not VPP even when using dpdk plugin), although multiple subsequent tests (even with VPP) may be affected. The real cause is probably in NIC firmware or driver, but CSIT can be better at detecting port status as a workaround.
  • test: testpmd (also l3fwd but hidden by CSIT-1896)
  • frequency: always (almost)
  • testbed: 3n-alt, 3n-snr
  • example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-dpdk-perf-mrr-weekly-master-3n-alt/65/log.html.gz#s1-s1-s1-s3-t1-k2-k4
  • ticket: CSIT-1848
  • note: The infra cause got fixed, but there still is CSIT-1904 with the same consequences.

(L) 3n-alt: Tests failing until 40Ge Interface comes up

(L) 1n-aws: TRex NDR PDR ALL IP4 scale and L2 scale tests failing with 50% packet loss

  • last update: 2023-06-21
  • work-to-fix: hard
  • rca: Perhaps AWS limits the number of usable IPv4 addresses.
  • test: ip4scale2m
  • frequency: always
  • testbed: 1n-aws
  • example: 1n-aws
  • ticket: CSIT-1876
  • note: The issue is still present, but the affected scales are no longer used in trending.