Jump to content

IBM RS/6000 Troubleshooting Guide: Difference between revisions

From RetroTechCollection
Remove red link to a page that does not exist
Rewrite from IBM's own code lists: most BIST/IPL/AIX code tables here were invented (e.g. 102 'ROM checksum', 110 'Power Good', 120-159 'subsystem' ranges, 551 'login prompt', SCSI '0540'). Now SA23-2631-06 App. C, 888 read-out procedure, 7013 500 power-failure codes (SA38-0531) and 7043 checkpoints (SA38-0512), all hosted
 
Line 1: Line 1:
This guide documents fault diagnosis for the '''[[IBM RS/6000]]''' family (machine types '''7011 / 7012 / 7013 / 7015 / 7020 / 7025 / 7026 / 7043 / 7044 / 7248'''). RS/6000 troubleshooting differs significantly from PC troubleshooting in that '''every RS/6000 has a 3- or 4-digit LED operator panel''' that displays both BIST (Built-In Self-Test) and POST codes, and that operating-system errors are reported with structured '''SRN''' (Service Request Number) codes that map directly to IBM-supplied FRU lists.<ref name="src1" /><ref name="src2" /><ref name="src3" /><ref name="src4" /><ref name="src5" /><ref name="src6" /><ref name="src7" /><ref name="src8" /><ref name="src9" />
<templatestyles src="Template:StyledTable/styles.css" />
[[File:IBM RS-6000 (photo).jpg|thumb|right|300px|IBM RS/6000]]
 
Fault diagnosis on the [[IBM RS/6000]] starts at the operator panel. The Micro Channel machines report their self-test, power-on and configuration progress, and any halt, as a number on a three-digit display; later PCI machines such as the 7043 use four-character firmware checkpoints and eight-character error codes. Hardware errors found by the diagnostics are reported as a service request number (SRN), which IBM's service documentation maps to field-replaceable units (FRUs). The code lists below are IBM's own; they differ between machine families, so use the service guide for the machine type in front of you.


== The Operator Panel ==
== Operator panel (Micro Channel machines) ==


Every RS/6000 carries an operator panel with a 3-digit or 4-digit LED display (some later 7026 / 7044 systems use a 16-character LCD instead). The display shows:
IBM's diagnostic operator guide describes these panel features:<ref name="og">IBM, ''POWERstation and POWERserver Diagnostic Programs: Operator Guide'', Version 2.1, SA23-2631-06, April 1992, pp. 1-5 to 1-7 ([[:File:IBM RS6000 Diagnostic Programs Operator Guide SA23-2631-6.pdf]]).</ref>


* '''Codes during power-on''' — BIST and POST progress; halt codes if a fault is detected.
* Power-on light: on when all voltages in the power supply are present and within limits and the fans are running. If a fan sensed by the supply stops or fails to start, the supply turns the system unit off.
* '''Codes after AIX boots''' — operating-system service codes; "888" flashing for kernel halt.
* Key mode switch: Secure prevents an initial program load (IPL) and disables the Reset button, and an IPL attempted in Secure shows 200. Normal loads the operating system from disk after POST and configuration. Service loads the diagnostic controller, searching diskette, then non-disk SCSI devices, then disks, then the network.
* '''Codes during shutdown''' — graceful shutdown progress.
* Reset button: restarts the system in the mode set by the key, reads out a crash or diagnostic message after a flashing 888, and starts a dump.
* Three-digit display: tracks the progress of the built-in self-test (BIST), power-on self-test (POST) and configuration programs; if a test stops, its number stays on the display.


The complete reference is the per-machine '''Operator Guide''' / '''Service Guide''' (SA38-05xx family) plus the kev009 mirror of IBM's "RS/6000 3-Digit Display Codes" document.<ref>http://ps-2.kev009.com/jasper/aix/rs-leds.errcodes.html</ref><ref>http://ps-2.kev009.com/pimpworks/ibm/aixled.html</ref>
Check stops, machine checks and other processor-detected errors reset the system and start an IPL. A second check stop during that IPL stops the system with 113; uncorrectable memory errors and addressing errors stop it with a flashing 888.<ref name="og" />


== BIST Codes (Hardware POST) ==
== Three-digit display codes ==


BIST runs first at power-on, before any firmware initialisation. BIST codes are the lowest-level fault indication.
The following are from Appendix C of IBM's 1992 operator guide.<ref name="ogc">IBM, SA23-2631-06, Appendix C, "Three-Digit Display Numbers", pp. C-1 to C-8.</ref> Numbers 100 to 199 belong to BIST, 200 to 299 to POST, and 500 to 999 to the configuration programs.


{| class="wikitable styled-table" style="width:100%; text-align:center;"
{| class="wikitable styled-table" style="width:100%; text-align:left;"
|+'''Common BIST / hardware POST codes'''
|+'''BIST indicators (selection)'''
! Code !! Meaning
! Code !! IBM's meaning
|-
| 100 || BIST is running. Cleared on success.
|-
| 101 || BIST is running. Cleared on success.
|-
| 102 || BIST checksum on the boot ROM failed. ROM corrupted or failed.
|-
|-
| 103 || BIST timed out. CPU or service-processor fault.
| 100 || BIST completed successfully; control passed to IPL ROS
|-
|-
| 104 || Equipment Check (general hardware fault).
| 101 || BIST started following reset
|-
|-
| 105 || ROS (Read-Only Storage) test failure.
| 102 || BIST started following power-on reset
|-
|-
| 106 || L2 cache failure (where fitted).
| 103 || BIST could not determine the system model number
|-
|-
| 110 || Power Good not received from PSU in time.
| 104 || Equipment conflict; BIST could not find the CBA
|-
|-
| 111 || Bus interface fault on the planar.
| 105 || BIST could not read from the OCS EPROM
|-
|-
| 112 || Watchdog timer expired during BIST.
| 106 || BIST detected a module error
|-
|-
| 113 || Reset issued by service processor — usually transient.
| 111 || OCS stopped; BIST detected a module error
|-
|-
| 120–129 || Memory subsystem BIST failures.
| 112 || Checkstop during BIST; results could not be logged out
|-
|-
| 130–139 || I/O subsystem BIST failures.
| 113 || BIST checkstop count greater than 1
|-
|-
| 140–149 || SCSI subsystem BIST failures.
| 121-127 || CRC checks of the OCS EPROM, NVRAM (OCS and time-of-day areas) and 8752 EPROM; odd numbers are "bad CRC" results
|-
|-
| 150–159 || Graphics adapter BIST failures.
| 888 || BIST did not start
|}
|}


The exact per-code FRU list is in each machine's service guide. The above is the cross-family pattern.
{| class="wikitable styled-table" style="width:100%; text-align:left;"
 
|+'''POST indicators (selection)'''
== IPL POST Codes (Firmware Boot) ==
! Code !! IBM's meaning
 
After BIST, the firmware (proprietary ROS on POWER1/POWER2; PReP on 7248; Open Firmware on CHRP) runs IPL. IPL codes are in the 200–400 range.
 
{| class="wikitable styled-table" style="width:100%; text-align:center;"
|+'''Common IPL / firmware POST codes'''
! Code !! Meaning
|-
| 200 || '''Mode switch in SECURE position''' — boot attempted but blocked.<ref>http://ps-2.kev009.com/jasper/aix/older.rs-leds.errcodes.html</ref>
|-
| 201 || '''Checkstop during IPL''' — major hardware fault, CPU halted.<ref>http://ps-2.kev009.com/jasper/aix/older.rs-leds.errcodes.html</ref>
|-
|-
| 202 || NVRAM read failure. Replace NVRAM module (see [[IBM RS/6000 Maintenance Guide]]).
| 200 || IPL attempted with the key in the Secure position
|-
|-
| 203 || NVRAM CRC failure. Re-enter SMS configuration.
| 201 || IPL ROM test failed or checkstop occurred (irrecoverable)
|-
|-
| 204 || NVRAM write failure.
| 211 || IPL ROM CRC comparison error (irrecoverable)
|-
|-
| 205 || Service processor not responding.
| 212 || RAM POST memory configuration error or no memory found (irrecoverable)
|-
|-
| 210 || Memory configuration error. Re-seat memory cards.
| 213 || RAM POST failure (irrecoverable)
|-
|-
| 211 || Memory ECC test failure.
| 214 || Power status register failed (irrecoverable)
|-
|-
| 220 || I/O slot configuration error.
| 215 || A low voltage condition is present (irrecoverable)
|-
|-
| 221 || PCI bus configuration error (CHRP machines).
| 217 || End of boot list encountered
|-
|-
| 222 || PCI device enumeration failure.
| 221 || NVRAM CRC error during IPL in Normal mode; reset NVRAM by doing an IPL in Service mode
|-
|-
| 230 || SCSI controller not responding.
| 222-237 || Attempting a Normal mode IPL or restart from the devices in the NVRAM or IPL ROM device lists (planar devices, SCSI, 9333, 7012 direct-bus disk, Ethernet, token ring)
|-
|-
| 231 || SCSI bus configuration error.
| 229 || Cannot IPL from any device in the NVRAM list, or the list has no valid entries
|-
|-
| 232 || No boot device found on SCSI.
| 242-258 || The same device searches in Service mode
|-
|-
| 240 || Boot disk not in boot list.
| 262 || No keyboard detected on the keyboard port
|-
|-
| 241 || Boot image corrupt.
| 281-293 || POST of the keyboard, parallel and serial ports, graphics, token ring, Ethernet, adapter slots, standard I/O, SCSI and 7012 direct-bus disk
|-
|-
| 250 || Network boot started.
| 290 || IOCC POST error (irrecoverable)
|-
|-
| 260 || Diagnostic mode requested (key in SERVICE position).
| 297 || System model number does not compare between OCS and ROS (irrecoverable)
|-
|-
| 299 || IPL completed; AIX kernel handed control.
| 299 || IPL ROM passed control to the loaded program code
|}
|}


The specific code-to-FRU mapping is in the per-machine service guide. The above is the cross-family pattern.
Configuration codes 500 to 508 show the configuration manager querying the native I/O slot and slots 1 to 8, 510 and 511 device configuration starting and completing, and 520 to 539 bus configuration and errors in the configuration manager or its ODM database (IBM marks 521 to 529, 531, 532, 534 and 536 irrecoverable). 551 means IPL varyon is running, 552 that it failed, and 553 that IPL phase 1 is complete. Numbers from 581 upwards show individual software components, devices and adapters being configured; if the display stops on one, that item is the first suspect.<ref name="ogc" />
 
== Flashing 888 ==


== AIX-Generated Codes (After Kernel Boot) ==
A flashing 888 means a crash message or a diagnostic message is waiting. To read it (IBM's procedure, pp. 3-14 and 3-15):<ref name="og888">IBM, SA23-2631-06, "Reading Flashing 888 Numbers", pp. 3-14 to 3-16, and Appendix C, p. C-8.</ref>


Once AIX is running, the LED panel is driven by the kernel. Codes in the 500–900 range indicate runtime / operational events.
# Have paper ready. Make sure the key is in Normal or Service. Hold Reset for about one second each time you press it.
# Press Reset once and record the number: this is the message type (102, 103 or 105).
# Type 102 (crash): press Reset and record the crash code, press again and record the dump status code. If 888 then flashes again, the message is complete; otherwise a 103 or 105 message follows.
# Type 103 or 105 (diagnostic): press Reset twice to read the six-digit SRN in two halves, then keep pressing to read up to four FRU location codes, each shown as a c0x prefix followed by eight numbers, until 888 flashes again.
# The system must be turned off to recover from the halt.


{| class="wikitable styled-table" style="width:100%; text-align:center;"
{| class="wikitable styled-table" style="width:100%; text-align:left;"
|+'''Common AIX runtime codes'''
|+'''Type 102 crash codes (Appendix C, p. C-8)'''
! Code !! Meaning
! Code !! IBM's meaning
|-
|-
| 500 || Init started.
| 000 || Unexpected system interrupt
|-
|-
| 511 || Filesystem checks running.
| 200-207 || Machine check: memory bus parity (200), memory timeout (201), memory card failure (202), address out of range (203), write to ROS (204), uncorrectable address parity (205), uncorrectable ECC error (206), unidentified (207)
|-
|-
| 517 || /etc/inittab being processed.
| 300 || Data storage interrupt from the processor
|-
|-
| 520 || Init complete, daemons starting.
| 32x, 38x || Data storage interrupt from an I/O exception on the IOCC or SLA (x is the bus unit ID)
|-
|-
| 540 || Network initialising.
| 400 || Instruction storage interrupt
|-
|-
| 551 || Login prompt presented.
| 500, 501 || External interrupt: scrub memory bus parity error, or unidentified error
|-
|-
| 553 || /etc/rc.tcpip running.
| 51x-54x || External interrupt on IOCC x: DMA memory bus error (51x), channel check (52x), bus timeout (53x), keyboard check (54x)
|-
|-
| 581 || ODM (Object Data Manager) reconfiguring.
| 558 || Not enough space to continue the program load
|-
|-
| 700 || Program Interrupt — kernel panic or invalid instruction.<ref>https://sysadminera.com/2017/02/04/the-led-codes-of-aix/</ref>
| 700 || Program interrupt
|-
|-
| 888 || '''Unexpected system halt''' (kernel panic / hardware fault). Code flashes. See below.<ref>https://sysadminera.com/2017/02/04/the-led-codes-of-aix/</ref>
| 800 || Floating point not available
|}
|}


=== The 888 Halt Sequence ===
== Service request numbers ==
 
When the kernel panics, the operator panel cycles through a sequence beginning with a flashing '''888'''. The format is:
 
  888 102 xxx yyy
 
Where:
 
* '''102''' — software-induced crash (the kernel called the panic routine).
* '''xxx''' — crash code:
:: '''300''' = Data Storage Interrupt (DSI; bad memory access).
:: '''700''' = Program Interrupt (invalid instruction, trap, panic).
:: '''0c0''' = successful kernel dump completed.
:: '''0c5''' = kernel dump attempted but failed.
* '''yyy''' — dump status code.


Press the '''Reset''' button on the operator panel to advance through the sequence. Record every code before resetting.
When the diagnostic programs find a problem they report an SRN, which the service documentation maps to the FRUs to replace, in order. Type 103 and 105 messages give the SRN and the location codes of up to four FRUs.<ref name="og888" /> The SRN-to-FRU lists are in IBM's diagnostic service guides and the machine service guides, not in the operator guide.


== SRN — Service Request Numbers ==
== 7013 500 series: power supply failure codes ==


When AIX detects a hardware error or when a diagnostic is run, the error is reported as an '''SRN''' (Service Request Number) in the format:
On the 7013 500 series, when the supply detects a power failure it turns the power light off and shows a status on the three-digit display. Blanks in all three positions mean no failure, or none that can be detected. A zero is not a failure.<ref name="sg531">IBM, ''RS/6000 7013 500 Series Installation and Service Guide'', SA38-0531-00, November 1996, MAP 1520, pp. 2-1520-1 and 2-1520-2, and "Power Supply Connectors", pp. 1-15 and 1-16 ([[:File:IBM RS6000 7013 500 Series Installation and Service Guide SA38-0531-00.pdf]]).</ref>


  sss-rrr
{| class="wikitable styled-table" style="width:90%; text-align:left;"
 
|+'''7013 500 series power failure status'''
Where '''sss''' is the source code (CPU, memory, SCSI, etc.) and '''rrr''' is the reason. SRNs map directly to the FRU list in the per-machine service guide. Run '''diag''' from AIX to see active and historical SRNs.
! Position !! Digit !! Meaning
 
Examples (from kev009 LED codes reference):<ref>http://ps-2.kev009.com/jasper/aix/older.rs-leds.errcodes.html</ref>
 
* '''101-xxx''' — Memory.
* '''165-xxx''' — Planar logic.
* '''201-xxx''' — Memory test.
* '''221-xxx''' — Memory adapter.
* '''260-xxx''' — Display adapter.
* '''651-xxx''' — Internal SCSI device.
* '''767-xxx''' — Graphics adapter.
 
== SCSI-Specific Codes ==
 
The 0540-class SCSI error codes appear during IPL or AIX boot when the SCSI subsystem cannot enumerate devices.
 
{| class="wikitable styled-table" style="width:100%; text-align:center;"
|+'''Common SCSI LED / SRN patterns'''
! Code area !! Meaning !! First action
|-
|-
| 0540 || SCSI controller initialisation timeout || Reseat SCSI controller card; check SCSI cable
| 1 || 8 / 4 / u || Fan on P49B failed / fan on P49C failed / both
|-
|-
| 0552 || SCSI bus parity error || Replace SCSI cable, then terminator, then drive
| 2 || 8 || Power supply fan on P47 failed
|-
|-
| 0556 || SCSI hard reset failure || Check for missing terminator at end of bus
| 2 || 4 || Fan on P46A or P46 failed
|-
|-
| 0581 || No bootable device on SCSI bus || Run SMS, verify boot list
| 2 || 2 || Fan on P49A failed
|}
 
== Stiction Diagnosis ==
 
If the operator panel halts at a 230 / 232 / 0540 / 0556 SCSI code immediately after power-on and the SCSI drive is silent (not spinning), suspect spindle stiction. See [[IBM RS/6000 Maintenance Guide]] for the field stiction fix.
 
== Mode Switch ==
 
The operator panel includes a key switch with positions:
 
* '''NORMAL''' — boot the default OS from the boot list.
* '''SERVICE''' — boot the standalone diagnostic image (from disk or CD-ROM).
* '''SECURE''' — block all boots (used for transport; produces a '''200''' code on attempted boot).
 
Always check the mode switch position before troubleshooting an apparent boot failure — a key turned to SECURE is a very common "first day in the lab" symptom.
 
== SMS — System Management Services (CHRP Machines) ==
 
On 7025 / 7026 / 7043 / 7044, the SMS menu is the user-facing configuration interface. Entry keys at POST:
 
* '''F1''' — graphical SMS (at the keyboard icon).
* '''F4''' — text-mode SMS (English).
* '''F5''' — boot from default boot list ('''5''' on an ASCII terminal).
* '''F6''' — boot from the multiboot menu.
* '''F8''' — Open Firmware "ok" prompt.
 
If the system halts at a POST code before reaching the keyboard icon, SMS cannot be entered — diagnose the halt first.
 
== Memory Faults ==
 
* '''201''' / '''211''' / '''221''' — memory test or memory adapter failure. Reseat memory cards; on POWER2 39H/397 the memory must be in matched pairs.
* On 7012-39H / 7012-397 / 7013-590 / 7013-595 — pulling one card of a pair will halt POST with a 210 / 211 code.
* ECC scrubbing on long-uptime systems can produce gradual "memory degraded" entries in '''errpt''' — replace the indicated DIMM/SIMM before the second bit error in a word fails outright.
 
== Graphics Faults ==
 
* '''No video, system POSTs to 299 / 551 (login prompt) on serial console''' — graphics adapter problem. Try a different GXT card.
* '''Video on initial POST but blank by login prompt''' — AIX driver loaded for a card that isn't present, or wrong DSP file. Boot to single-user mode and run '''lsdev -C''' to verify the graphics device.
* '''GXT800P / GXT3000P pulling down +5 V''' — tantalum decoupling cap failed short. Recap.
 
== NVRAM Faults ==
 
NVRAM-class faults appear as:
 
* '''202 / 203 / 204''' at IPL — NVRAM read / CRC / write failure.
* '''Boot list reset to defaults every power cycle''' — Dallas TimeKeeper battery dead.
* '''Clock at epoch (1 Jan 1970 or 1 Jan 1980) after every cold boot''' — same.
 
Fix: replace the Dallas TimeKeeper module (see [[IBM RS/6000 Maintenance Guide]]).
 
== Service Processor Faults (Later Machines) ==
 
7025 / 7026 / 7044 carry a separate service processor that runs independently of the main CPU. If the service processor fails:
 
* '''No LED panel display at power-on''' — service processor not booting.
* '''LED display stuck at "STBY" or "0000"''' — service processor running but not transitioning to IPL.
* '''Service processor log full''' — accumulated thermal / fan / ECC events; clear via '''cfgmgr''' or the SP menu.
 
Service processor reset on 7025-F50 / F80 is via the small recessed switch on the rear of the chassis.
 
== PSU Faults ==
 
* '''Dead — no fans, no power''': PSU input rectifier or bulk capacitor. Discharge before any work.
* '''Fans spin briefly, then click-retry''': Power Good not asserted. Could be PSU fold-back, shorted planar tantalum, or '''on 7013 500-series''', a fan sense signal missing.<ref>https://www.ardent-tool.com/RS6000/docs/pdf/38053100.pdf</ref>
* '''Boots cold, fails when warm''': aged secondary electrolytics.
* '''Audible whine, smell of fish''': RIFA X2 cap venting.
* '''Rails low/high''': PSU feedback path issue. Recap.
 
== Diagnostic Workflow ==
 
# Confirm the mode switch is in NORMAL.
# Power on. Watch the LED panel.
# If LED stops at a code, look up the code in the per-machine service guide first; if not found, in the cross-family table above; finally in kev009.<ref>http://ps-2.kev009.com/jasper/aix/rs-leds.errcodes.html</ref>
# If LED reaches 299 but AIX does not start, switch to a serial console (DB-9 on the system, 9600/8/N/1) and watch for kernel boot messages.
# If AIX boots but '''errpt''' shows ongoing hardware events, run '''diag''' to get the SRN.
# Cross-reference the SRN to the FRU list in the service guide; replace the indicated FRU(s).
 
== Common Faults and Resolutions ==
 
* '''No POST, no LED activity''' — PSU dead. Check rails. Check internal fuse if fitted.
* '''Halts at 200''' — mode switch in SECURE. Turn key.
* '''Halts at 201''' — checkstop. Reseat CPU MCM card; on POWER1/POWER2 7012/7013, suspect SMD electrolyte leakage on planar.
* '''Halts at 202 / 203''' — NVRAM dead. Replace Dallas module.
* '''Halts at 211''' — memory pair mismatch on POWER2. Reseat memory cards.
* '''Halts at 230 / 232 / 0540 / 0556''' — SCSI fault. Cable, terminator, drive stiction.
* '''Halts at 260''' — diagnostic mode (NORMAL key was in SERVICE).
* '''POST reaches 299 but AIX dies during boot''' — corrupted AIX rootvg. Boot from AIX install CD and use the maintenance shell.
* '''Boot loops to SMS''' — boot list points at a missing device; correct via SMS → Multiboot.
* '''Random reboots under load''' — PSU rails sagging, or thermal event due to failed CPU fan.
* '''Graphics card not detected''' — Open Firmware did not enumerate it; reseat and check for tantalum failure.
 
== Service Processor Logs (CHRP Machines) ==
 
On 7025 / 7026 / 7044, the service processor maintains an event log accessible via:
 
* The SP menu (boot to SMS, then "Service Processor Setup" → "Read Service Processor Log").
* AIX command '''snap -a''' (collects all logs to /tmp/ibmsupt for IBM service).
* AIX command '''errpt -a''' (formatted error report).
 
== Reading the 3-digit LED codes ==
 
The RS/6000 shows its boot and diagnostic progress, and any fault, as a '''3-digit code on the operator-panel LED display'''. Read the code first — it points straight at the failing area:
 
<templatestyles src="Template:StyledTable/styles.css" />
[[File:IBM RS-6000 (photo).jpg|thumb|right|300px|IBM RS/6000. Source: Wikimedia Commons.]]
{| class="wikitable styled-table" style="width:80%; text-align:left;"
|+'''RS/6000 operator-panel codes (selection)'''
! Code !! Meaning
|-
|-
| 215 || Low-voltage condition (irrecoverable) — power-supply fault
| 2 || 6, c, u, inverted F || Combinations of 4+2, 8+2, 8+4 and 8+4+2
|-
|-
| 221 || NVRAM CRC error — reset by re-running IPL in Service mode
| 3 || 8 || Temperature in the power supply excessive
|-
|-
| 223 || Attempting to IPL from the SCSI devices in the NVRAM boot list (stuck here = boot-device/SCSI problem)
| 3 || 4 || Power failure inside the power supply
|-
|-
| 21c || L2 cache not detected
| 3 || 2 || Power failure outside the power supply
|-
|-
| F57 || Bad or low backup battery
| 3 || 1 || Loss of primary power
|-
| F6A / F7A / F7B || SCSI init / NVRAM init / NVRAM CRC check
|}
|}


The full published code list is in the references.<ref name="rs6k">[http://ps-2.kev009.com/jasper/aix/rs-leds.errcodes.html RS/6000 3-Digit Display Codes], Kev009; and the IBM RS/6000 service manuals. Source for the operator-panel LED codes and the NVRAM-battery behaviour.</ref>
If several failures are shown and one is in the right-hand position, IBM says to correct that one first. Any fan connector (P46A, P46B, P49A, P49B, P49C) without a fan must have the fan sense jumper fitted: the supply will not run with any of them open.<ref name="sg531" />


== NVRAM battery ==
== 7043 43P: firmware checkpoints and error codes ==


The RS/6000 keeps its configuration and IPL device list in battery-backed NVRAM. A '''dead or removed battery''' loses the settings and can leave the machine unbootable or "confused" — some 43P systems are hard to recover after the backup battery is pulled. Replace a low battery (F57), and after any NVRAM loss re-enter the boot configuration in Service mode.<ref name="rs6k" />
The 7043 Models 140, 150 and 240 use four-character firmware checkpoints on the operator panel (E100 onwards) and eight-character error codes. IBM's service guide lists them with repair actions:<ref name="sg512">IBM, ''RS/6000 7043 43P Series Service Guide'', SA38-0512-03, October 1998, chapters 2 to 4 and 7 ([[:File:IBM RS6000 7043 43P Series Service Guide SA38-0512-03.pdf]]).</ref>


== SCSI and memory ==
* A system that stops at a memory checkpoint (E122, E213, E214, E220 or E3xx) goes to the memory problem MAP.
* A display alternating between E1FD and another Exxx code: record both.
* 20E00004, "battery drained or needs replacement": replace the battery, then the system planar. 28030005 is an RTC battery error. IBM notes that NVRAM and real-time clock errors can be caused by low battery voltage.
* When the keyboard indicator appears during start-up, F1 (or 1 on an ASCII terminal) starts System Management Services, whose Utilities menu holds the error log. F5 (or 5) boots the diagnostics from the default list, for example from the diagnostic CD-ROM.


Most RS/6000 boot hangs are the '''SCSI subsystem''' (the codes step through SCSI init and the IPL-device search) or bad '''memory SIMMs'''. Reseat the SIMMs and the SCSI cabling and terminators, and confirm the boot device against the NVRAM IPL list.<ref name="rs6k" />
== A diagnostic sequence ==


== References ==
# Check the key mode switch: 200 on the display means it is in Secure.
# Watch the display from power-on and note where it stops. A number that stays up identifies the unsuccessful test or the device being configured.
# If the power light goes out and a status appears (7013 500 series), use the power failure table above; check that unused fan connectors have their sense jumpers.
# For a flashing 888, read out the whole message before turning the system off.
# Run the diagnostics in Service mode and record the SRN and location codes, then look them up in the service guide for the machine type.
 
== Related pages ==


<references>
* [[IBM RS/6000]], [[IBM RS/6000 Maintenance Guide]], [[IBM RS/6000 Capacitor Replacement Guide]]
<ref name="src1">[http://ps-2.kev009.com/jasper/aix/rs-leds.errcodes.html RS/6000 3-Digit Display Codes — kev009 mirror]. Authoritative LED code reference.</ref>
<ref name="src2">[http://ps-2.kev009.com/jasper/aix/older.rs-leds.errcodes.html RS/6000 older 3-Digit Display Codes]. Specific BIST/IPL code listings.</ref>
<ref name="src3">[http://ps-2.kev009.com/pimpworks/ibm/aixled.html IBM AIX LED diagnostic codes — kev009]. Cross-reference list.</ref>
<ref name="src4">[https://sysadminera.com/2017/02/04/the-led-codes-of-aix/ The LED Codes of AIX — sysadminera]. Modern restated reference for 888 halt sequence.</ref>
<ref name="src5">[https://www.ardent-tool.com/RS6000/docs/pdf/38053100.pdf IBM SA38-0531-00 — RS/6000 7013 500-series Installation and Service Guide]. PSU + planar service.</ref>
<ref name="src6">[https://sharktastica.co.uk/resources/docs/IBM_SA38-0512-03_RS6000_98_4.pdf IBM SA38-0512-03 — RS/6000 7043 / 7248 Service Guide]. CHRP / PReP service.</ref>
<ref name="src7">[https://www.manualslib.com/manual/901093/Ibm-Rs-6000-Enterprise-Server-M80.html?page=356 IBM RS/6000 M80 Service Manual]. Common firmware error codes chapter (applies broadly across CHRP RS/6000s).</ref>
<ref name="src8">[https://www.infania.net/misc/rs6000-firmware/7043140I.html 7043-140 firmware and SMS keys]. F1/F4/F5/F6/F8 entry-key reference.</ref>
<ref name="src9">[http://ps-2.kev009.com/tl/techlib/qna/faxes/html/krn/krn4.htm IBM TechLib krn4 — RS/6000 boot keys].</ref>
</references>


== Related Pages ==
== References ==


* [[IBM RS/6000]]
<references />
* [[IBM RS/6000 Maintenance Guide]]
* [[IBM RS/6000 Capacitor Replacement Guide]]


{{Navbox-IBMComputers|state=collapsed}}
{{Navbox-IBMComputers|state=collapsed}}

Latest revision as of 22:00, 30 September 2026

IBM RS/6000

Fault diagnosis on the IBM RS/6000 starts at the operator panel. The Micro Channel machines report their self-test, power-on and configuration progress, and any halt, as a number on a three-digit display; later PCI machines such as the 7043 use four-character firmware checkpoints and eight-character error codes. Hardware errors found by the diagnostics are reported as a service request number (SRN), which IBM's service documentation maps to field-replaceable units (FRUs). The code lists below are IBM's own; they differ between machine families, so use the service guide for the machine type in front of you.

Operator panel (Micro Channel machines)

[edit | edit source]

IBM's diagnostic operator guide describes these panel features:[1]

  • Power-on light: on when all voltages in the power supply are present and within limits and the fans are running. If a fan sensed by the supply stops or fails to start, the supply turns the system unit off.
  • Key mode switch: Secure prevents an initial program load (IPL) and disables the Reset button, and an IPL attempted in Secure shows 200. Normal loads the operating system from disk after POST and configuration. Service loads the diagnostic controller, searching diskette, then non-disk SCSI devices, then disks, then the network.
  • Reset button: restarts the system in the mode set by the key, reads out a crash or diagnostic message after a flashing 888, and starts a dump.
  • Three-digit display: tracks the progress of the built-in self-test (BIST), power-on self-test (POST) and configuration programs; if a test stops, its number stays on the display.

Check stops, machine checks and other processor-detected errors reset the system and start an IPL. A second check stop during that IPL stops the system with 113; uncorrectable memory errors and addressing errors stop it with a flashing 888.[1]

Three-digit display codes

[edit | edit source]

The following are from Appendix C of IBM's 1992 operator guide.[2] Numbers 100 to 199 belong to BIST, 200 to 299 to POST, and 500 to 999 to the configuration programs.

BIST indicators (selection)
Code IBM's meaning
100 BIST completed successfully; control passed to IPL ROS
101 BIST started following reset
102 BIST started following power-on reset
103 BIST could not determine the system model number
104 Equipment conflict; BIST could not find the CBA
105 BIST could not read from the OCS EPROM
106 BIST detected a module error
111 OCS stopped; BIST detected a module error
112 Checkstop during BIST; results could not be logged out
113 BIST checkstop count greater than 1
121-127 CRC checks of the OCS EPROM, NVRAM (OCS and time-of-day areas) and 8752 EPROM; odd numbers are "bad CRC" results
888 BIST did not start
POST indicators (selection)
Code IBM's meaning
200 IPL attempted with the key in the Secure position
201 IPL ROM test failed or checkstop occurred (irrecoverable)
211 IPL ROM CRC comparison error (irrecoverable)
212 RAM POST memory configuration error or no memory found (irrecoverable)
213 RAM POST failure (irrecoverable)
214 Power status register failed (irrecoverable)
215 A low voltage condition is present (irrecoverable)
217 End of boot list encountered
221 NVRAM CRC error during IPL in Normal mode; reset NVRAM by doing an IPL in Service mode
222-237 Attempting a Normal mode IPL or restart from the devices in the NVRAM or IPL ROM device lists (planar devices, SCSI, 9333, 7012 direct-bus disk, Ethernet, token ring)
229 Cannot IPL from any device in the NVRAM list, or the list has no valid entries
242-258 The same device searches in Service mode
262 No keyboard detected on the keyboard port
281-293 POST of the keyboard, parallel and serial ports, graphics, token ring, Ethernet, adapter slots, standard I/O, SCSI and 7012 direct-bus disk
290 IOCC POST error (irrecoverable)
297 System model number does not compare between OCS and ROS (irrecoverable)
299 IPL ROM passed control to the loaded program code

Configuration codes 500 to 508 show the configuration manager querying the native I/O slot and slots 1 to 8, 510 and 511 device configuration starting and completing, and 520 to 539 bus configuration and errors in the configuration manager or its ODM database (IBM marks 521 to 529, 531, 532, 534 and 536 irrecoverable). 551 means IPL varyon is running, 552 that it failed, and 553 that IPL phase 1 is complete. Numbers from 581 upwards show individual software components, devices and adapters being configured; if the display stops on one, that item is the first suspect.[2]

Flashing 888

[edit | edit source]

A flashing 888 means a crash message or a diagnostic message is waiting. To read it (IBM's procedure, pp. 3-14 and 3-15):[3]

  1. Have paper ready. Make sure the key is in Normal or Service. Hold Reset for about one second each time you press it.
  2. Press Reset once and record the number: this is the message type (102, 103 or 105).
  3. Type 102 (crash): press Reset and record the crash code, press again and record the dump status code. If 888 then flashes again, the message is complete; otherwise a 103 or 105 message follows.
  4. Type 103 or 105 (diagnostic): press Reset twice to read the six-digit SRN in two halves, then keep pressing to read up to four FRU location codes, each shown as a c0x prefix followed by eight numbers, until 888 flashes again.
  5. The system must be turned off to recover from the halt.
Type 102 crash codes (Appendix C, p. C-8)
Code IBM's meaning
000 Unexpected system interrupt
200-207 Machine check: memory bus parity (200), memory timeout (201), memory card failure (202), address out of range (203), write to ROS (204), uncorrectable address parity (205), uncorrectable ECC error (206), unidentified (207)
300 Data storage interrupt from the processor
32x, 38x Data storage interrupt from an I/O exception on the IOCC or SLA (x is the bus unit ID)
400 Instruction storage interrupt
500, 501 External interrupt: scrub memory bus parity error, or unidentified error
51x-54x External interrupt on IOCC x: DMA memory bus error (51x), channel check (52x), bus timeout (53x), keyboard check (54x)
558 Not enough space to continue the program load
700 Program interrupt
800 Floating point not available

Service request numbers

[edit | edit source]

When the diagnostic programs find a problem they report an SRN, which the service documentation maps to the FRUs to replace, in order. Type 103 and 105 messages give the SRN and the location codes of up to four FRUs.[3] The SRN-to-FRU lists are in IBM's diagnostic service guides and the machine service guides, not in the operator guide.

7013 500 series: power supply failure codes

[edit | edit source]

On the 7013 500 series, when the supply detects a power failure it turns the power light off and shows a status on the three-digit display. Blanks in all three positions mean no failure, or none that can be detected. A zero is not a failure.[4]

7013 500 series power failure status
Position Digit Meaning
1 8 / 4 / u Fan on P49B failed / fan on P49C failed / both
2 8 Power supply fan on P47 failed
2 4 Fan on P46A or P46 failed
2 2 Fan on P49A failed
2 6, c, u, inverted F Combinations of 4+2, 8+2, 8+4 and 8+4+2
3 8 Temperature in the power supply excessive
3 4 Power failure inside the power supply
3 2 Power failure outside the power supply
3 1 Loss of primary power

If several failures are shown and one is in the right-hand position, IBM says to correct that one first. Any fan connector (P46A, P46B, P49A, P49B, P49C) without a fan must have the fan sense jumper fitted: the supply will not run with any of them open.[4]

7043 43P: firmware checkpoints and error codes

[edit | edit source]

The 7043 Models 140, 150 and 240 use four-character firmware checkpoints on the operator panel (E100 onwards) and eight-character error codes. IBM's service guide lists them with repair actions:[5]

  • A system that stops at a memory checkpoint (E122, E213, E214, E220 or E3xx) goes to the memory problem MAP.
  • A display alternating between E1FD and another Exxx code: record both.
  • 20E00004, "battery drained or needs replacement": replace the battery, then the system planar. 28030005 is an RTC battery error. IBM notes that NVRAM and real-time clock errors can be caused by low battery voltage.
  • When the keyboard indicator appears during start-up, F1 (or 1 on an ASCII terminal) starts System Management Services, whose Utilities menu holds the error log. F5 (or 5) boots the diagnostics from the default list, for example from the diagnostic CD-ROM.

A diagnostic sequence

[edit | edit source]
  1. Check the key mode switch: 200 on the display means it is in Secure.
  2. Watch the display from power-on and note where it stops. A number that stays up identifies the unsuccessful test or the device being configured.
  3. If the power light goes out and a status appears (7013 500 series), use the power failure table above; check that unused fan connectors have their sense jumpers.
  4. For a flashing 888, read out the whole message before turning the system off.
  5. Run the diagnostics in Service mode and record the SRN and location codes, then look them up in the service guide for the machine type.
[edit | edit source]

References

[edit | edit source]
  1. ↑ 1.0 1.1 IBM, POWERstation and POWERserver Diagnostic Programs: Operator Guide, Version 2.1, SA23-2631-06, April 1992, pp. 1-5 to 1-7 (File:IBM RS6000 Diagnostic Programs Operator Guide SA23-2631-6.pdf).
  2. ↑ 2.0 2.1 IBM, SA23-2631-06, Appendix C, "Three-Digit Display Numbers", pp. C-1 to C-8.
  3. ↑ 3.0 3.1 IBM, SA23-2631-06, "Reading Flashing 888 Numbers", pp. 3-14 to 3-16, and Appendix C, p. C-8.
  4. ↑ 4.0 4.1 IBM, RS/6000 7013 500 Series Installation and Service Guide, SA38-0531-00, November 1996, MAP 1520, pp. 2-1520-1 and 2-1520-2, and "Power Supply Connectors", pp. 1-15 and 1-16 (File:IBM RS6000 7013 500 Series Installation and Service Guide SA38-0531-00.pdf).
  5. ↑ IBM, RS/6000 7043 43P Series Service Guide, SA38-0512-03, October 1998, chapters 2 to 4 and 7 (File:IBM RS6000 7043 43P Series Service Guide SA38-0512-03.pdf).