Difference between revisions of "CSIT/TestFailuresTracking"
From fd.io
								< CSIT
												
				|  (Move CSIT-1911 to Past.) |  (CSIT-1904 not visible on 2n-clx due to jobspecs.) | ||
| Line 126: | Line 126: | ||
| * ticket: [https://jira.fd.io/browse/CSIT-1906 CSIT-1906] | * ticket: [https://jira.fd.io/browse/CSIT-1906 CSIT-1906] | ||
| * note: Currently not visible in trending as CSIT-1905 hits first. | * note: Currently not visible in trending as CSIT-1905 hits first. | ||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| ==== (L) 3n-icx: negative ipackets on TB38 AVF 4c l2patch ==== | ==== (L) 3n-icx: negative ipackets on TB38 AVF 4c l2patch ==== | ||
| Line 221: | Line 210: | ||
| * example: [https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-report-coverage-2210-2n-icx/10/log.html.gz#s1-s1-s1-s1-s1 2n-icx] | * example: [https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-report-coverage-2210-2n-icx/10/log.html.gz#s1-s1-s1-s1-s1 2n-icx] | ||
| * ticket: [https://jira.fd.io/browse/CSIT-1800 CSIT-1800] | * ticket: [https://jira.fd.io/browse/CSIT-1800 CSIT-1800] | ||
| + | |||
| + | ==== (M) 2n-clx: DPDK 23.03 link failures ==== | ||
| + | |||
| + | * last update: 2023-06-21 | ||
| + | * work-to-fix: medium | ||
| + | * rca: No link comes up on some NICs, investigation continues. | ||
| + | * test: Testpmd (no vpp). Around half of tested+NIC combinations are affected. | ||
| + | * frequency: always since 23.03.0 got released | ||
| + | * testbed: multiple | ||
| + | * example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-dpdk-perf-mrr-weekly-master-2n-clx/189/log.html.gz#s1-s1-s1-s1-t1-k2-k4 | ||
| + | * ticket: [https://jira.fd.io/browse/CSIT-1904 CSIT-1904] | ||
| + | * note: No longer affects 2n-clx trending just because the affected NICs were removed from jobspecs. | ||
| ==== (L) 2n-clx, 2n-icx: nat44ed cps 16M sessions scale fail ==== | ==== (L) 2n-clx, 2n-icx: nat44ed cps 16M sessions scale fail ==== | ||
Revision as of 10:32, 21 June 2023
Contents
- 1 CSIT Test Failure Clasification
- 2 CSIT Test Fixing Priorities
- 3 Current Failures
- 3.1 Deterministic Failures
- 3.1.1 In Trending
- 3.1.1.1 (H) AVF suite setup fails if previous suite was also AVF
- 3.1.1.2 (M) 2n-icx: interface down in nginx tests
- 3.1.1.3 (M) hoststack: ip4udpscale1cl10s-ldpreload-iperf3 times out waiting for strace
- 3.1.1.4 (M) 3n-snr: All hwasync wireguard tests failing when trying to verify device
- 3.1.1.5 (M) 1n-aws: TRex mlrsearch fails to find NDR & PDR due to AWS rate limiting (5min total test duration)
- 3.1.1.6 (M) 3n-alt, 3n-snr: testpmd no traffic forwarded
- 3.1.1.7 (M) 3n-alt: Tests failing until 40Ge Interface comes up
- 3.1.1.8 (M) 2n-spr 200Ge2P1Cx7Veat: TRex sees port line rate as 100 Gbps
- 3.1.1.9 (M) 2n-spr: zero traffic on cx7 rdma
- 3.1.1.10 (L) 3n-icx: negative ipackets on TB38 AVF 4c l2patch
 
- 3.1.2 Not In Trending
- 3.1.2.1 (M) all testbeds: some vpp 9000B tests
- 3.1.2.1.1 (M) tests with 9000B payload frames not forwarded over vhost interfaces
- 3.1.2.1.2 tests with 9000B payload frames not forwarded over memif interfaces
- 3.1.2.1.3 9000B payload frames not forwarded over tunnels due to violating supported Max Frame Size (VxLAN, LISP, SRv6)
- 3.1.2.1.4 (M) 9000b all AVF tests are failing to forward traffic
 
- 3.1.2.2 (M) 2n-clx, 2n-icx, 2n-zn2: DPDK testpmd 9000b tests on xxv710 nic are failing with no traffic
- 3.1.2.3 (M) 2n-clx, 2n-icx: all Geneve tests with 1024 tunnels fail
- 3.1.2.4 (M) 2n-clx: DPDK 23.03 link failures
- 3.1.2.5 (L) 2n-clx, 2n-icx: nat44ed cps 16M sessions scale fail
- 3.1.2.6 (L) 2n-clx, 2n-icx: nat44det imix 1M sessions fails to create sessions
 
- 3.1.2.1 (M) all testbeds: some vpp 9000B tests
 
- 3.1.1 In Trending
- 3.2 Occasional Failures
- 3.3 Rare Failures
- 3.3.1 In Trending
- 3.3.1.1 (M) 3n-icx, 3n-snr: 1518B IPsec packets not passing
- 3.3.1.2 (M) all testbeds: mlrsearch fails to find NDR rate
- 3.3.1.3 (M) all testbeds: AF_XDP mlrsearch fails to find NDR rate
- 3.3.1.4 (L) all testbeds: vpp create avf interface failure in multi-core configs
- 3.3.1.5 (L) all testbeds: nat44det 4M and 16M scale 1 session not established
 
 
- 3.3.1 In Trending
 
- 3.1 Deterministic Failures
- 4 Past Failures
CSIT Test Failure Clasification
All known CSIT failures grouped and listed in the following order:
- Always failing followed by sometimes failing.
-  Always failing tests:
- Most common use cases followed by less common.
 
-  Sometimes failing tests:
-  Most frequently failing followed by less frequently failing. 
- High frequency 50%-100%
- medium frequency 10%-50%
- low frequency 0%-10%.
 
- Within each sub-group: most common use cases followed by less common.
 
-  Most frequently failing followed by less frequently failing. 
CSIT Test Fixing Priorities
Test fixing work priorities defined as follows:
- (H)igh priority, most common use cases and most common test code.
- (M)edium priority, specific HW and pervasive test code issue.
- (L)ow priority, corner cases and external dependencies.
Current Failures
Deterministic Failures
In Trending
(H) AVF suite setup fails if previous suite was also AVF
- last update: 2023-05-15
- work-to-fix: low
- rca: After a recent change, CSIT attempt to bind an already bound driver.
- test: All AVF suites if the previous suite running on the NIC was also AVF.
- frequency: always since 2023-05-09
- testbed: all (if having AVF supported NIC)
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-3n-icx/265/log.html.gz#s1-s1-s1-s3-s2-k1-k5-k3-k1-k1-k1-k1-k1-k1-k1-k1-k1-k1-k1
- ticket: CSIT-1913
- note: proposed fix https://gerrit.fd.io/r/c/csit/+/38831
(M) 2n-icx: interface down in nginx tests
- last update: 2023-05-03
- work-to-fix: medium
- rca: Likely an infra issue for TG with AB on NIC using ICE.
- test: All nginx tests except for xxv710 with dpdk driver.
- frequency: always since 2023-04-28
- testbed: 2n-icx
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-hoststack-daily-master-2n-icx/51/log.html.gz#s1-s1-s1-s1-s1-t1-k2-k8-k3-k2-k1
- ticket: CSIT-1910
(M) hoststack: ip4udpscale1cl10s-ldpreload-iperf3 times out waiting for strace
- last update: 2023-04-19
- work-to-fix: medium
- rca:
- test: Only the two ip4udpscale1cl10s-ldpreload-iperf3 tests
- frequency: always since 2023-03-04
- testbed: 3n-icx
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-hoststack-daily-master-3n-icx/41/log.html.gz#s1-s1-s1-s1-s8-t1-k2-k5-k12
- ticket: CSIT-1908
(M) 3n-snr: All hwasync wireguard tests failing when trying to verify device
- last update: before 2023-01-31
- work-to-fix: hard
- rca: Missing QAT driver. Symptom: Failed to bind PCI device 0000:f4:00.0 to c4xxx on host 10.30.51.93
- test: hwasync wireguard
- frequency: always
- testbed: 3n-snr
- example: 3n-snr
- ticket: CSIT-1883
(M) 1n-aws: TRex mlrsearch fails to find NDR & PDR due to AWS rate limiting (5min total test duration)
- last update: 2023-02-09
- work-to-fix: hard
- rca:
- test: ip4scale2m
- frequency: always
- testbed: 1n-aws
- example: 1n-aws
- ticket: CSIT-1876
- note: The root cause can be shared environment in aws cloud. We may need to use a smaller scale there.
(M) 3n-alt, 3n-snr: testpmd no traffic forwarded
- last update: 2023-02-09
- work-to-fix: medium
- rca: DUT-DUT link takes too long to come up on some testbeds. This happens *after* a test case with a DPDK app (not VPP even when using dpdk plugin), although multiple subsequent tests (even with VPP) may be affected. The real cause is probably in NIC firmware or driver, but CSIT can be better at detecting port status as a workaround.
- test: testpmd (also l3fwd but hidden by CSIT-1896)
- frequency: always (almost)
- testbed: 3n-alt, 3n-snr
- example: 3n-alt, 3n-snr, 3n-snr
- ticket: CSIT-1848
(M) 3n-alt: Tests failing until 40Ge Interface comes up
- last update: 2023-02-09
- work-to-fix: medium
- rca: DUT-DUT link takes too long to come up due to CSIT-1848.
- test: first tests in order
- frequency: always (almost, depends on run order)
- testbed: 3n-alt (3n-snr link does not take that long)
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-3n-alt/155/log.html.gz#s1-s1-s1-s1-s1-t1
- ticket: CSIT-1890
(M) 2n-spr 200Ge2P1Cx7Veat: TRex sees port line rate as 100 Gbps
- last update: 2023-04-19
- work-to-fix: medium
- rca: TRex has hard cap on perceived line rate.
- test: All tests on 2n-spr with 200Ge2P1Cx7Veat NIC.
- frequency: always (since 2n-spr was set up)
- testbed: 2n-spr
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-2n-spr/15/log.html.gz#s1-s1-s1-s5-s19-t1-k2-k9-k9-k10-k1-k1-k1-k11
- ticket: CSIT-1905
- note: Fix will be in TRex v3.03, possible workarounds being discussed.
(M) 2n-spr: zero traffic on cx7 rdma
- last update: 2023-04-19
- work-to-fix: medium
- rca: VPP reports "tx completion errors", more investigation ongoing.
- test: All tests on 2n-spr with 200Ge2P1Cx7Veat NIC and RDMA driver.
- frequency: always (since 2n-spr was set up)
- testbed: 2n-spr
- ticket: CSIT-1906
- note: Currently not visible in trending as CSIT-1905 hits first.
(L) 3n-icx: negative ipackets on TB38 AVF 4c l2patch
- last update: 2023-04-19
- work-to-fix:
- rca:
- test: TB38 AVF 4c e810cq, earlier l2patch nowadays eth-l2xcbase
- frequency: always
- testbed: 3n-icx (only TB38, never TB37)
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-3n-icx/214/log.html.gz#s1-s1-s1-s5-s3-t3-k2-k9-k8-k13-k1-k2
- ticket: CSIT-1901
Not In Trending
(M) all testbeds: some vpp 9000B tests
- last update: 2023-02-09
- work-to-fix: hard
- rca: VPP code: 34839: dpdk: cleanup MTU handling. CSIT needs to rework how it sets MTU / max frame rate (CSIT-1797). Some tests will continue failing due to missing support on VPP side, we will open specific Jira tickets for those.
- test: see sub-items
- frequency: always
- testbed: all
- examples: see sub-items
- ticket: CSIT-1809
- gerrit: https://gerrit.fd.io/r/c/csit/+/37824
(M) tests with 9000B payload frames not forwarded over vhost interfaces
- last update: 2023-02-09
- work-to-fix: hard
- test: 9000B + vhostuser
- testbed: 2n-skx, 3n-skx, 2n-clx
- examples: 3n-skx vhostuser
- ticket: CSIT-1809
tests with 9000B payload frames not forwarded over memif interfaces
- last update: 2023-02-09
- work-to-fix: hard
- test: 9000B + memif
- testbed: 2n-skx, 3n-skx, 2n-clx
- examples: 2n-skx Memif
- ticket: CSIT-1808
9000B payload frames not forwarded over tunnels due to violating supported Max Frame Size (VxLAN, LISP, SRv6)
- last update: 2023-02-09
- work-to-fix: medium
- test: 9000B + (IP4 tunnels VXLAN, IP4 tunnels LISP, Srv6, IpSec)
- testbed: 2n-icx, 3n-icx
- examples: 2n-icx VXLAN, 3n-icx
- ticket: CSIT-1801
(M) 9000b all AVF tests are failing to forward traffic
- last update: 2023-02-09
- work-to-fix: hard
- test: 9000B + AVF
- testbed: 3n-icx
- examples: 3n-icx ip4base
- ticket: CSIT-1885
(M) 2n-clx, 2n-icx, 2n-zn2: DPDK testpmd 9000b tests on xxv710 nic are failing with no traffic
- last update: 2023-02-09
- work-to-fix: medium
- rca: The DPDK app only attempts to set MTU once, but if interface is down (CSIT-1848) it fails. As a workaround, MTU could be set on Linux interface before starting the DPDK app.
- test: DPDK testpmd 9000b
- frequency: always
- testbed: 2n-clx, 2n-icx, 2n-zn2
- example: 2n-clx, 2n-icx
- ticket: CSIT-1870
- note: Vratko will fix, either in general workaround for CSIT-1848 or in a separate change.
(M) 2n-clx, 2n-icx: all Geneve tests with 1024 tunnels fail
- last update: before 2023-01-31
- work-to-fix: hard
- rca: VPP crash, Failed to add IP neighbor on interface geneve_tunnel258
- test: avf-ethip4--ethip4udpgeneve-1024tun-ip4base 64B 1518B IMIX 1c 2c 4c
- frequency: always
- testbed: 2n-skx, 2n-clx, 2n-icx
- example: 2n-icx
- ticket: CSIT-1800
(M) 2n-clx: DPDK 23.03 link failures
- last update: 2023-06-21
- work-to-fix: medium
- rca: No link comes up on some NICs, investigation continues.
- test: Testpmd (no vpp). Around half of tested+NIC combinations are affected.
- frequency: always since 23.03.0 got released
- testbed: multiple
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-dpdk-perf-mrr-weekly-master-2n-clx/189/log.html.gz#s1-s1-s1-s1-t1-k2-k4
- ticket: CSIT-1904
- note: No longer affects 2n-clx trending just because the affected NICs were removed from jobspecs.
(L) 2n-clx, 2n-icx: nat44ed cps 16M sessions scale fail
- last update: before 2023-01-31
- work-to-fix: hard
- rca: VPP crash, Failed to set NAT44 address range on host 10.30.51.44 (connections-per-second tests only)
- test: 64B-avf-ethip4tcp-nat44ed-h262144-p63-s16515072-cps-ndrpdr 1c 2c 4c, 64B-avf-ethip4udp-nat44ed-h262144-p63-s16515072-cps-ndrpdr 1c 2c 4c
- frequency: always
- testbeds: 2n-skx, 2n-clx, 2n-icx
- example: 2n-icx, 2n-clx
- ticket: CSIT-1799
(L) 2n-clx, 2n-icx: nat44det imix 1M sessions fails to create sessions
- last update: before 2023-01-31
- work-to-fix: hard
- rca:
- test: IMIX over 1M sessions bidir
- frequency: always
- testbed: 2n-skx, 2n-clx, 2n-icx
- example: 2n-icx
- ticket: CSIT-1884
Occasional Failures
In Trending
(H) 2n-icx: NFV density VPP does not start in container
- last update: before 2023-01-31
- work-to-fix: hard
- rca:
- test: all subsequent
- frequency: medium
- testbed: 2n-icx
- example: 2n-icx mrr, 2n-icx ndrpdr
- ticket: CSIT-1881
- note: Once VPP breaks, all subsequent tests fail. Even all subsequent builds will be failing until Peter makes TB working again. Although it's failing with medium frequency when it happens it breaks all subsequent builds on the TB therefore [H] priority.
(M) 2n-clx: e810 mlrsearch tests packets forwarding in one direction
- last update: before 2023-01-31
- work-to-fix: hard
- rca:
- test: e810Cq ip4base, ip6base
- frequency: high
- testbed: 2n-clx
- example: 2n-clx
- ticket: CSIT-1864
(M) 3n-icx, 3n-snr: wireguard 100 and 1000 tunnels mlrsearch tests failing with 2c and 4c
- last update: before 2023-01-31
- work-to-fix: easy
- rca:
- test: wireguard 100 tunnels and more
- frequency: high
- testbed: 3n-icx, 3n-snr
- examples: 3n-icx
- ticket: CSIT-1886
(M) 3n-tsh: vpp in VM starting too slowly
- last update: before 2023-02-22
- work-to-fix: medium
- rca: perhaps related to numa, investigation continues
- test: 3n-tsh: sporadic VM vhost
- frequency: high
- testbed: 3n-tsh
- example: 3n-tsh, 3n-tsh
- ticket: CSIT-1877
Rare Failures
In Trending
(M) 3n-icx, 3n-snr: 1518B IPsec packets not passing
- last update: before 2023-01-31
- work-to-fix: hard
- rca:
- test: all AVF crypto
- frequency: low
- testbed: 3n-skx, 3n-icx, 3n-snr
- example: 3n-icx daily, 3n-snr, 3n-icx weekly
- ticket: CSIT-1827
(M) all testbeds: mlrsearch fails to find NDR rate
- last update: before 2023-04-19
- work-to-fix: hard
- rca: One (not sure whether only) possible symptom is ierrors on TRex side. Not sure it is TRex error or VPP sending mangled packets.
- test: Crypto, Ip4, L2, Srv6, Vm Vhost (all packet sizes, all core configurations affected)
- frequency: low
- testbed: 3n-tsh, 3n-alt, 2n-clx
- example: 2n-icx
- ticket: CSIT-1804
(M) all testbeds: AF_XDP mlrsearch fails to find NDR rate
- last update: before 2023-01-31
- work-to-fix: hard
- rca:
- test: af-xdp multicore tests
- frequency: low
- testbed: 2n-clx, 2n-skx, 2n-tx2, 2n-icx
- example: 2n-skx, 2n-clx
- ticket: CSIT-1802
- note: This is mainly observed in iterative and coverage. It's very low frequency ~ 1 out of 100
(L) all testbeds: vpp create avf interface failure in multi-core configs
- last update: 2023-02-06
- work-to-fix: hard
- rca: issue in Intel FVL driver
- test: multicore AVF
- frequency: low
- testbed: all testbeds
- example: 2n-clx, 3n-icx
- ticket: CSIT-1782
- note: A long standing issue without a final permanent fix.
(L) all testbeds: nat44det 4M and 16M scale 1 session not established
- last update: 2023-02-14
- work-to-fix: hard
- rca: unknown
- test: nat44det udp 4m and 16m (64k is ok, 1m can fail but rarely than bigger scales)
- frequency: low
- testbed: 2n-zn2, 2n-skx, 2n-icx, 2n-clx
- example: 2n-zn2, 2n-clx
- ticket: CSIT-1795
Past Failures
(M) 2n-clx: 100Ge2P1Cx556A not recognized in mlx5-core tests
- last update: 2023-05-15
- work-to-fix: medium
- rca:
- test: All with 100Ge2P1Cx556A and mlx5-core. At least TB27 is affected.
- frequency: always since 2023-04-28
- testbed: 2n-clx
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-2n-clx/1317/log.html.gz#s1-s1-s1-s5-s7-t1-k2-k6-k3-k1-k1-k1-k1-k1-k1-k3-k1-k1
- ticket: CSIT-1909
(L) 3n-alt: hugepage leak
- last update: 2023-05-02
- work-to-fix: medium
- rca:
- test: All vhost tests when number of free hugepages gets low
- frequency: happened only once (2023-04-24) so far, but multiple runs were affected until huge pages were freed manually
- testbed: 3n-alt
- example: https://s3-logs.fd.io/vex-yul-rot-jenkins-1/csit-vpp-perf-mrr-daily-master-3n-alt/210/log.html.gz#s1-s1-s1-s6-s1-t1-k2-k9-k6-k1
- ticket: CSIT-1911
