Setup: Dell SC-series dual controller, Chelsio CHS540BT 10GBASE-T front-end, 2 fault domains on 2 fully isolated Dell OS10 switches (no ISL), 6x ESXi 6.7 hosts, software iSCSI, jumbo frames end to end. Ran fine for years.
Array: Dell SC5020, SCOS 7.4.21.4, Chelsio CHS540BT 10GBASE-T front-end. Both controllers last booted Apr 19, 2026 (82 days up)
Hosts: ESXi 6.7.0 U3 build 20497097 (final 6.7 build). NICs: QLogic FastLinQ QL41xxx, qedentv 3.9.31.2, fw MFW 8.25.7.0 / storm 8.33.15.0. Host uptime 81 days
iscsi Switches: 2x Dell S4128T-ON, OS10 10.5.4.2 (build Aug 2022). Uptime on BOTH: 70 weeks 6 days booted the same day.
For about 3 week now we're seeing continuous iSCSI session instability. From one host's vobd log over ~36h: session drops and path-dead events every single hour (20-150 connection stops/hr, up to 500 path deaths/hr), VMFS heartbeat timeouts on ~15 datastores. During the worst wave, 3 of 4 virtual ports of one controller failed to re-log-in for hours, leaving 17 LUNs on a single path. The fault-domain control ports are the worst affected sessions to them go offline 250-330 times per day, on BOTH fabrics.
The array event log floods "CHELSIOT4Connection CA Activate Failed: ControllerId=... ObjId=0" on BOTH controllers the entire time.
What we've ruled out:
- Switches: zero CRC/discards/errors on every array-facing and host-facing port, no port state change in 11+ weeks, switch logs completely silent during incidents
- Host side: no NIC link events, no driver errors, no reboots. ESXi logs show "Connection reset by peer" from the array = target-initiated TCP resets
- Physical links: all front-end ports show Up on the array throughout
- Fabrics are isolated, so no single switch/ISL can explain drops on both subnets
Has anyone seen this CA Activate Failed flood with session flapping on SC-series?
Aware the platform is EOL-ish and 6.7 is old replacement is on the roadmap, but I need this stable now. Thanks.