== PS1 IPP Czar Logs for the week 2010.01.24 - 2010.01.30 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2011.01.24 === * eam : warp seemed to be slow: no progress for about 1 hour. the book was full of DONE runs. I stopped processing, ran 'process_cleanup warpPendingSkyCell' to clear the book, and restarted processing. It ran fine after that * eam : burntool seemed to have gotten stuck. I looked on the new burntool state ippMonitor page and found one of the imfiles did not seem to be making progress. Looking at the pantasks (control status), I noticed that the job for that cell had been running for a very long time (>500 sec). I went to the machine where it was running and noticed that it was hanging on access /data/ipp033.0 (which crashed over the weekend). I used force.umount to clear the mount point, and things moved along from there * bills 15:53 Removed ipp053 from distribution host lists and restarted distribution pantasks. It seemed sluggish anyways. * bills 16:00 cleared some magicDSRun revert faults. I need to automate this! * bills 16:09 updated magic_destreak_cleanup.pl to *not* delete the original uncensored diff stage cmf files. * bills 16:10 set label STS.201009 back to active. * bills 16:25 Executed stacktool -updaterun -set_state drop -stack_id 216256 -set_note 'fails due to problem in ticket 1427' * bills 19:37 several faults have appeared setting STS.201009 back to inactive * bills 21:19 still lots of faults. I suspect that the rsync processes running on this node are related. I did neb-host --host ipp008 --state repair === Tuesday : 2011.01.25 === Bill is czar today * 04:21 Many many faults. ipp012 filesystem is read-only from many nodes. ssh is rejected. Stopping all processing for a few minutes to investigate. * 04:25 ipp012 console unresponsive. 'Only output from console: INIT: Id "s0" r[220286.000299] Kernel panic - not syncing: Attempted to kill init!' cycling power * 04:37 processing restarted . Space is running out. * 04:39 There is a very large postage stamp request (> 15000 jobs from MPIA) Setting MPIA labels to inactive for now. * 04:55 distribution pantasks jobs for ipp012 are all failing. setting ipp012 to off. * 07:11 many nodes over 98%. Stopped processing except for summit copy and registration. All of last night's data has been copied but there are 308 still working on registration & burntool. * 07:23 ippdb00:/tmp is full which is causing nebulous errors. Ran: 'sudo mv /tmp/nebulous_server.log /export/ippdb00.0/nebulous_server.log-20110125' Still has zero free. * 08:00 fixed register exp problem being caused by invalid value for $default_host in registration pantasks.Changed it from ipp023 to any * 08:30 burntool is proceeding. * 08:30 since magic is way backlogged and since it doesn't use much space I've turned distribution back to run with destreak.off * 09:26 fixed broken magic_cleanup script. Data is being recovered on ipp053. Turning processing back on. * 09:41 turned on STS label set priority to 1000. Once those exposures are done and distributed we can clean up the data. * 11:13 turned off chip to allow the sts warps to progress faster. * 11:00 ipp053 is < 98% now. Queued magicRuns on ipp050 for cleanup. It's down to 97% as well. * 13:15 chip.on === Wednesday : 2011.01.26 === === Thursday : 2011.01.27 === === Friday : 2011.01.28 === === Saturday : 2011.01.29 === === Sunday : 2011.01.30 ===