== PS1 IPP Czar Logs for the week 2014.12.01 - 2014.12.07 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2014.12.01 === * 09:53 Bill: restarted distribution pantasks which was sluggish with diff stage enabled. SAS distribution is all finished except for diff stage. * 13:00 Bill: ~ippsky is rerunning staticsky and skycal for sas.37 * 17:00 MEH: setting up ippx049+ for ippsky staticsky/cal to use -- manually adding when ready, ~3x should be fine for 50G ram * ippx049--x088 have been setup and added temporarily to a x2 group (need split by pdu?), after Haydn's email x089-x092 also -- see x2 group for ones still down * 22:00 MEH: nightly processing not doing well w/ ippmd and ippsky running (or something else?) -- ipp083-087 overloaded/not responding and stalling registration === Tuesday : 2014.12.02 === * 07:11 Bill: started up fullforce in ~ippsky/staticsky (staticsky and skycal are finished except for one skycell that faulted with a data error. It is running now * 07:30 MEH: registration backed up quite a bit, too many nodes in use -- ippmd stop until nightly finishes * '''looks like something may have started ~0200, nightly rate dropped greatly then as well.. and not just for stdsci''' -- looking at the storage node network, appears the red storage nodes hit a data block about that time and there it is... * 08:30 MEH: taking some xNNN nodes out of ippsky until nightly finishes.. then probably could do 4x on those for FF * 08:40 MEH: an exposure (o6993g0619o) is stalled in registration.. been a while so restart summitcopy and registration (then stdsci once nightly is finished) -- had to manually run {{{ regtool -dbname gpc1 -revertprocessedexp -exp_id 829654 }}} * 09:30 MEH: ipp056,058,060 have space and could be moved out of repair to help ease load on new storage nodes and get nightly finished * same for ipp012,013,029,030 * now get to see those load >50 -- so even when get 10G for ipp083-088 setup, the other 1G systems will get large loads? * 12:00 MEH: warp growth curve fault -- {{{ warptool -dbname gpc1 -updateskyfile -fault 0 -set_quality 42 -warp_id 1241149 -skycell_id skycell.1616.069 }}} * 12:30 MEH: nightly finally finished, while waiting for storage node reboots, going to do the regular restart of stdsci and set to stop * 12:50 MEH: adding host groups s4,s5 for the new storage hosts ipp067-070,ipp071-088 respectively. will start testing running jobs on them including deep stacks and should not be allocated for any other jobs at this time === Wednesday : 2014.12.03 === === Thursday : 2014.12.04 === === Friday : 2014.12.05 === === Saturday : 2014.12.06 === * '''revised''' -- planned downtime MRTC-B 590B Lipoa Automatic Transfer Switch (ATS) -- Saturday December 6th. They will be starting our shutdown processes at 9am and will hopefully start turning things back on at 3pm. * '''all processing stopped before ~7am''' * Haydn on site to start power down of systems @7am === Sunday : 2014.12.07 ===