Jump to content

DECsystem-10 Troubleshooting Guide

From RetroTechCollection
DECsystem-10 KI10 processor cabinets (Gah4, CC BY-SA 4.0)

Fault-finding on a DECsystem-10 from DEC's field service documents: the KL10 Maintenance Guide (EK-OKL10-MG-003), the KL10-C power system description (EK-PWR20-SD-001), the MH10 Maintenance Manual (EK-MH10-MM-003) and the 1971 DECsystem-10 Technical Summary.[1][2][3][4] Power and cooling faults come first, because the power controls shut the whole system down for them; then memory; then the processor through its console front end and diagnostics.

DEC's method

[edit | edit source]

DEC's KL10 guide teaches a seven-step fault analysis, fed by three ways of gathering information (asking the operator, observing, and operating the machine):[5]

  1. State the problem: what is happening that should not, what is not happening that should, and under what conditions.
  2. Form a hypothesis of the most probable cause.
  3. Design an experiment or test.
  4. Predict the result.
  5. Carry out the experiment.
  6. Evaluate the result.
  7. Correct the fault, or repeat steps 1 to 7.

DEC then confirms the repair either by putting the fault back or by running for a time window long enough to prove it gone.[5]

The system will not power up, or shuts down

[edit | edit source]

KL10 processor

[edit | edit source]

The 863 power control shuts the system down for any of eight faults, lighting the FAULT lamp on the switch panel and an individual LED on the 863: insufficient air flow at any of four sensors, over-temperature (an unused input on the KL10-C), a tripped breaker to the CPU supply, or either of two cooling-assembly doors open. After a fault the system stays down for at least 13 seconds.[6]

KL10-C power faults (from DEC's power system description)
Symptom Likely cause Check
FAULT lamp lit, 863 air-flow LED lit Fan stopped, blocked filter or failed air-flow sensor Fans, filters, sensor; the H770 +15 V supply feeds the sensors
FAULT lamp lit, 863 door LED lit Cooling-assembly door open Close and latch the doors
FAULT lamp lit, CPU breaker LED lit Breaker on the H761 series pass assembly tripped Find which breaker has tripped in the H761
Powers up but CPU logic dead; G8010 or G8011 LEDs off Failed regulator card in the H761 DEC: apart from tripped breakers, LEDs off are the only sign of a failed regulator card
Will not power up again straight after power-off Restart lock-out Wait 13 seconds
POWER lamp blinking Override switch in the 863 is on Turn it off once the fault is cleared

[6][7]

A power failure is detected when the line falls below 95 V on a 115 V system (190 V on 230 V) or leaves 47–63 Hz. The 863 asserts POWER WARNING, the KL10 sees the power line flag (bit 30 of CONI APR) and shuts down in order, and the PDP-11/40 front end traps to location 24 with 2 ms to save its state. When power returns the front end reloads the KL10 microcode, which is lost in a power failure. DEC notes that only AC power faults restart automatically; every other fault must be dealt with by hand.[8]

KA10 and KI10

[edit | edit source]

DEC's power-fail circuit on these processors interrupts the program so the monitor can save its registers. The KI10 monitors all three phases and can restart automatically when power returns. Temperature sensors in the equipment shut the power down on overheating, which also triggers the power-fail interrupt.[9]

Memory cabinets

[edit | edit source]

An MH10 cabinet powers itself down if any of its four air-flow switches trips (the OVERTEMP lamp lights) or a stack or logic door panel is opened with power on. It will not power up again for four seconds after power-off.[10] If the cabinet does not respond to the system power switch, check that its 857 REMOTE/LOCAL switch is at REMOTE, that its input-voltage selector matches the supply, and that the three-wire power control bus cables are connected.[11]

Regulators

[edit | edit source]

An H744 +5 V regulator limits its output current at about 30–35 A and fires a crowbar SCR across the output above about +6.0 V; an H754 limits at about 10 A and has a crowbar on each output. A crowbarred output stays at 0 V until power is cycled.[10] Check each rail at the backplane against DEC's limits in DECsystem-10 Power Supply Guide.[12]

Memory faults

[edit | edit source]

Parity errors

[edit | edit source]

With the MH10 maintenance panel's OVERRIDE/CHECK switch at CHECK, a parity error freezes the data, address, active port and read/write request lights for the failing reference while the memory carries on.[13][14] DEC's guidance:[13]

  • Write parity errors usually come from the processor (CPU or data channel) or the memory bus cable to the port shown as active; a bad transceiver or data register in the MH10's port control logic can also cause them.
  • Read parity errors come from the core memory and control logic.
  • Each of the two controllers has its own parity checker, and which controller is active depends on the interleave mode, so changing the interleave mode and watching which controller reports helps isolate the data path.

Margins and stack sets

[edit | edit source]

Run the memory diagnostic MAINDEC-10-DDMMG with the STRB, THRESH and CUR margin switches, high and low, one at a time; a memory that fails only at margin is marginal.[15][14] The G236, G116 and H224 modules of a 32K stack set are factory-matched; DEC replaces all three together, and does not adjust the G236's strobe or drive-current jumpers in the field.[15] A failing 64K bank can be switched out with its bank select switch so the system can run on the remaining memory.[16]

KL10 processor faults

[edit | edit source]

Gathering error information

[edit | edit source]

Before gathering KL10 error information, disable automatic reload so the front end does not restart the machine over the evidence. DEC's procedure for TOPS-20 and TOPS-10 release 603 or later is to enter the front end's PARSER (Ctrl+\), type SET CONSOLE PROGRAMMER, then SET NO RELOAD (PARSER replies RELOAD ENABLE: OFF), then QUIT. On TOPS-10, DEC's Procedure 2 also sets bit 03 of the monitor word DEBUGF with FILDDT.[17]

DEC's Procedure 3 then saves the machine state through the front end:[17]

  1. In PARSER, SET CON MAINT for maintenance mode, then EXAMINE DTE-20 for the DTE20 status.
  2. Load and start the KLDCP diagnostic console program (MCR BOO, then DBOOT).
  3. In KLDCP, examine the DTE status and diagnostic registers (EE 174434, EE 174430), reset the DTE, stop the clock (FX 0), print all CRAM and registers (ALL), read the APR (FR 110) and all diagnostic functions (FR 100,177).
  4. Step the clock to find the PC several times, and check with EC whether the current CRAM location is the microcode's halt loop. If it is not, the KL10 is hung in an unknown state.
  5. Examine accumulators 16 and 17; if that fails, follow DEC's jumper procedure in the guide. Then read APR, PAG, PI and DTE status, the page-fail word, PC and flags, internal and external memory status, each RH20 Massbus controller and each internal channel.

The guide follows these procedures with fault isolation tables for page-fail traps and hardware faults, separately for internal and external memory.[18]

Remote diagnosis

[edit | edit source]

DEC's KL10 systems have the KLINIK link, through which DEC's diagnostic centre could run diagnostics on a customer's machine remotely.[19] Volume II of DEC's KL10 maintenance guide covers the maintenance software, including KLDCP and DIAMON; it and a KLAD10 diagnostic "cookbook" are on bitsavers.[20][21]

[edit | edit source]

References

[edit | edit source]
  1. ↑ Digital Equipment Corporation, KL10 Maintenance Guide, Volume I, EK-OKL10-MG-003, April 1985 (File:KL10 Maintenance Guide Volume 1 (EK-OKL10-MG-003).pdf).
  2. ↑ Digital Equipment Corporation, DECSYSTEM-20 Power System, System Description, EK-PWR20-SD-001, April 1976 (File:KL10-C Power System Description (EK-PWR20-SD-001).pdf).
  3. ↑ Digital Equipment Corporation, MH10 Maintenance Manual, EK-MH10-MM-003, August 1977 (File:MH10 Core Memory Maintenance Manual (EK-MH10-MM-003).pdf).
  4. ↑ Digital Equipment Corporation, DECsystem-10 Technical Summary, 1971 (File:DECsystem-10 Technical Summary (1971).pdf).
  5. ↑ 5.0 5.1 DEC, EK-OKL10-MG-003, general information p. 44, "Seven Steps of Fault Analysis".
  6. ↑ 6.0 6.1 DEC, EK-PWR20-SD-001, pp. PWR20/2-4 to 2-10. 863 fault conditions, override, switch panel.
  7. ↑ DEC, EK-PWR20-SD-001, pp. PWR20/2-3 to 2-4. H761 regulator card LEDs and breakers; H770 for the air-flow sensors.
  8. ↑ DEC, EK-PWR20-SD-001, pp. PWR20/2-10 to 2-12. Power-failure protection and automatic restart.
  9. ↑ DEC, DECsystem-10 Technical Summary, 1971, p. 49. Safety features.
  10. ↑ 10.0 10.1 DEC, EK-MH10-MM-003, pp. 4-35 to 4-39. Air-flow and door interlocks, recycle delay, H744 and H754 protection.
  11. ↑ DEC, EK-MH10-MM-003, p. 2-3. 857 interconnection, REMOTE/LOCAL and line voltage switches.
  12. ↑ DEC, EK-OKL10-MG-003, general information p. 10. General power supply specifications.
  13. ↑ 13.0 13.1 DEC, EK-MH10-MM-003, pp. 5-3 to 5-4, section 5.4.2. Parity errors.
  14. ↑ 14.0 14.1 DEC, EK-MH10-MM-003, p. 3-6, Table 3-4. Maintenance panel switches.
  15. ↑ 15.0 15.1 DEC, EK-MH10-MM-003, pp. 5-1 to 5-4. Margins and stack set replacement.
  16. ↑ DEC, EK-MH10-MM-003, p. 1-2. Bank select switches.
  17. ↑ 17.0 17.1 DEC, EK-OKL10-MG-003, general information pp. 46–48, "Gathering KL10 System Error Information", Procedures 1 to 3.
  18. ↑ DEC, EK-OKL10-MG-003, general information, contents. Page fail trap and hardware fault isolation tables.
  19. ↑ DEC, EK-OKL10-MG-003, general information, "Important Telephone Numbers". Digital Diagnostic Center and KLINIK links.
  20. ↑ DEC, EK-OKL10-MG-003, "To the Reader", organisation of volumes I and II.
  21. ↑ Directory listing, bitsavers.org/pdf/dec/pdp10/KL10/, retrieved 2 October 2026.