* 2009-09-15T14:35:36 NFS issues; ls: cannot access /data/ipp051.0: Input/output error ; rebooted by gavin * 2009-09-18T16:32:55 Degraded unit: unit=0, port=18 (replaced) * 2010-01-05T16:20:30system unresponsive, ganglia [attachment:ipp029-loadgraph-2010-01-05.png load], [attachment:ipp029-memorygraph-2010-01-05.png memory], [attachment:ipp029-cpugraph-2010-01-05.png cpu], power cycled by gavin * 2011-5-12 5:37:16 Controller#1(PCI) IDE Channel # 9 Reading Error * 2011-07-06T11:26:14 system unresponsive, nothing on console, power cycled by gavin * 2011-07-07T03:28:44 system unresponsive, nothing on console, power cycled by gavin (2011-07-07T08:38:30) * 2011-10-03T23:48:00 system crash, console information shows the same kind of error as on ipp026: {{{ ipp029 login: [7653866.118158] [7653866.118158] HARDWARE ERROR [7653866.118158] CPU 5: Machine Check Exception: 4 Bank 0: b200004000000800 [7653866.118158] TSC 56ca1938fc43b0 [7653866.118158] This is not a software problem! [7653866.118158] Run through mcelog --ascii to decode and contact your hardware vendor ^^^^^^^^^^TRANSLATION^^^^^^^^^^^ (ipp029:~) cindy% /usr/sbin/mcelog --k8 --ascii < myerror HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 5 BANK 0 TSC 56ca1938fc43b0 MCG status:MCIP MCi status: Uncorrected error Error enabled Processor context corrupt MCA:BUS Level-0 Originated-request Generic Memory-access Request-timeout Error Model: STATUS b200004000000800 MCGSTATUS 4 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [7653866.118158] [7653866.118158] HARDWARE ERROR [7653866.118158] CPU 7: Machine Check Exception: 5 Bank 0: b200004000000800 [7653866.118158] RIP !INEXACT! 10: {mwait_idle+0x41/0x44} [7653866.118158] TSC 56ca1938fc43a0 [7653866.118158] This is not a software problem! [7653866.118158] Run through mcelog --ascii to decode and contact your hardware vendor [7653866.118158] ^^^^^^^^^^TRANSLATION^^^^^^^^^^^ (ipp029:~) cindy% /usr/sbin/mcelog --p4 --ascii < myerror2 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 7 BANK 0 TSC 56ca1938fc43a0 MCG status:RIPV MCIP MCi status: Uncorrected error Error enabled Processor context corrupt MCA:BUS Level-0 Originated-request Generic Memory-access Request-timeout Error Model: STATUS b200004000000800 MCGSTATUS 5 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [7653866.118158] HARDWARE ERROR [7653866.118158] CPU 7: Machine Check Exception: 5 Bank 5: b200120020080400 [7653866.118158] RIP !INEXACT! 10: {mwait_idle+0x41/0x44} [7653866.118158] TSC 56ca1938fc4fd8 [7653866.118158] This is not a software problem! [7653866.118158] Run through mcelog --ascii to decode and contact your hardware vendor [7653866.118158] ^^^^^^^^^^TRANSLATION^^^^^^^^^^^ (ipp029:~) cindy% /usr/sbin/mcelog --p4 --ascii < myerror3 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 7 BANK 5 TSC 56ca1938fc4fd8 MCG status:RIPV MCIP MCi status: Uncorrected error Error enabled Processor context corrupt MCA:Internal Timer error STATUS b200120020080400 MCGSTATUS 5 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [7653866.118158] HARDWARE ERROR [7653866.118158] CPU 5: Machine Check Exception: 4 Bank 5: b200120014040400 [7653866.118158] TSC 56ca1938fc5098 [7653866.118158] This is not a software problem! [7653866.118158] Run through mcelog --ascii to decode and contact your hardware vendor ^^^^^^^^^^TRANSLATION^^^^^^^^^^^ (ipp029:~) cindy% /usr/sbin/mcelog --p4 --ascii < myerror4 HARDWARE ERROR. This is *NOT* a software problem! Please contact your hardware vendor CPU 5 BANK 5 TSC 56ca1938fc5098 MCG status:MCIP MCi status: Uncorrected error Error enabled Processor context corrupt MCA:Internal Timer error STATUS b200120014040400 MCGSTATUS 4 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [7653866.118158] Kernel panic - not syncing: Machine check [7653866.118158] ------------[ cut here ]------------ [7653866.118158] WARNING: at kernel/smp.c:333 smp_call_function_mask+0x37/0x1d7() [7653866.118158] Modules linked in: coretemp w83627hf w83793 hwmon_vid k8temp autofs4 i2c_i801 i2c_core iTCO_wdt e1000e tg3 libphy e1000 xfs dm_snapshot dm_mirror dm_region_hash dm_log aacraid 3w_9xxx 3w_xxxx atp870u arcmsr aic7xxx scsi_wait_scan [7653866.118158] Pid: 26053, comm: ppSub Tainted: G M W 2.6.28-rc7-00105-gfeaf384 #4 [7653866.118158] Call Trace: [7653866.118158] <#MC> [] warn_on_slowpath+0x51/0x6d [7653866.118158] [] notify_update+0x2b/0x30 [7653866.118158] [] smp_call_function_mask+0x37/0x1d7 [7653866.118158] [] crash_kexec+0x17/0xef [7653866.118158] [] crash_kexec+0xe6/0xef [7653866.118158] [] native_smp_send_stop+0x1a/0x26 [7653866.118158] [] panic+0x95/0x13f [7653866.118158] [] release_console_sem+0x3e/0x1a5 [7653866.118158] [] release_console_sem+0x3e/0x1a5 [7653866.118158] [] release_console_sem+0x3e/0x1a5 [7653866.118158] [] __atomic_notifier_call_chain+0x74/0x83 [7653866.118158] [] __atomic_notifier_call_chain+0x0/0x83 [7653866.118158] [] mce_log+0x0/0x7f [7653866.118158] [] do_machine_check+0x2d5/0x378 [7653866.118158] [] machine_check+0x7f/0x90 [7653866.118158] <> <4>---[ end trace 4eaa2a86a8e2da22 ]--- }}} * 2011-10-06T10:00:00 system unresponsive, [wiki:Ipp029-crash-20111006 ipp029-hw-error], power cycled by gavin * 2011-10-08T12:41:06 system unresponsive, [wiki:Ipp029-crash-20111008 ipp029-hw-error], power cycled by gavin * 2011-10-09T15:38:56 system unresponsive, [wiki:Ipp029-crash-20111009 ipp029-hw-error], power cycled by gavin * 2011-10-11T13:26:16 system unresponsive, [wiki:Ipp029-crash-20111011 ipp029-hw-error], power cycled by gavin * 2011-10-11T14:11:16 system unresponsive, [wiki:Ipp029-crash-20111011T141116 ipp029-hw-error], power cycled by gavin * 2011-10-11T14:47:16 system unresponsive, [wiki:Ipp029-crash-20111011T144716 ipp029-hw-error], power cycled by gavin