== PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2013.04.15 === mark is czar * 07:10 MEH: nightly downloaded (before the morning shutdown again) and processed -- tweak_ssdiff to run early before any system work needs to be done today * 09:00 Bill: postage stamp pantasks was running very slowly (pcontrol spin) restarted it. Also restarted update. * 10:25 Bill: mysqld on ippc17 is not running. * MEH: ippc17 having several port/disk error messages.. * 10:40 MEH: stsci01 crashed over weekend even with neb-host repair. was rebooted into new 3.7.6 kernel, will put neb-host up as seems to be running okay like stsci06 -- stsci01 was one of the stsci machines that had not crashed since the power failure 4/2 and in neb-host repair, will probably want to start sequentially reboot all stsci machine soon into the new kernel.. * 10:50 MEH: shutting down majority of processing for ippdb work * 13:00 MEH: Serge et al finished ippdb swap and cleanup of apache node logs, system back up. ipp home directory <1G space left.. people need to clean home dir. moving ~ipp pantasks logs to archive and compressing. * 16:30 Gavin/Rita report that ippc17 looks to have experienced a motherboard failure, will need to be looked at again tomorrow. Datastore and postage server (as it uses the datastore to distribute images) will likely be down for 24 hours. * 17:00 Serge: Commented out ippc19 related backup on ipp@ipp001: /home/panstarrs/ipp/mysql-dump/ops_dump.csh * 17:10 Bill set rcserver.off as it uses the datastore, but distribution can be left running to do the preparation work. * 17:40 MEH: removing pstamp and publishing from roboczar until ippc17/datastore back * 23:50 MEH: looks like ipp025 has given up around 23:14.. SMP oops.. cannot log in, power cycled.. [wiki:ipp025_log] === Tuesday : 2013.04.16 === mark is czar * 06:17 Bill: postage stamp and update pantasks have been moved to ippc14 and started up. * 06:30 MEH: looks like registration may have had hangups over the early morning again.. may not make it through all processing before network down ~@0900. MD done so tweak_ssdiff * 8:13 Bill: set rcDestination.dbhost = 'ippc19' and restarted distribution pantasks. * 09:50 MEH: all processing stopped/shutdown including czartools for Rita/Gavin to swap switch connecting ippc01-ippc16. * 10:50 Haydn replaced BBU for raid in ipp064 * 10:55 Haydn adding 32G memory to ippc13 * lots to add.. * 17:50 MEH: everything stalled/faulting.. nebulous not working.. restarting apache ippc08,09 seemed to clear up. * 18:30 MEH: MOPS reporting still missing 3 exposures but nothing showing up. once system gets going again will try and trace back * 22:30 MEH: processing has been behaving like the parking brake is on.. many red 20TB disks and seen similar behavior before when this happened, but seems more extreme. stsci neb-host up no help. * even cleanup is painfully slow, so maybe an issue with nebulous or raid speed on a large portion of machines -- allowing a few more commented out nebservers (ippc04,06) but leaving ippc07 out due to network issue * cleared old ippc17.0 stalled mounts * stsci has dqstats process that uses datastore, had problem with missing dir so set to off and emailed Bill * removing ThreePi.WS.nightlyscience label to focus on primary nightly science === Wednesday : 2013.04.17 === Bill is czar today. Oh joy. * 06:40 processing is still painfully slow. camera_exp processing is taking > 4000 s. The ones I checked were all done with the real work and ever so slowly running the {{{ foreach file (@outputs { check fits and replicate } }}} * The problem is either network throughput or a nebulous bottleneck. Clearly having so many nodes working is not helping. I'm going to turn off a bunch and see if that improves * In email Gene suggests shutting everything down and restarting the nebulous apache servers. stopping stdscience * 07:36 changed my mind. I'm going to wait until I get to the office ~8:10 to restart things * 08:55 Bill stopped all processing. cleanup was stopped a couple of hours ago yet still has jobs running. Many blocked threads in nebulous mysql processlist. command = 'commit' * 09:20 Serge stopped the nebulous apaches. nebulous mysql threadlist emptied. Restarted apaches * 09:36 Bill restarted registration, pstamp, and stdscience pantasks. Will start others shortly. * 10:11 registration finished soon after restart. All pantasks are up except cleanup, deepstack, detrend, and replication * 10:11 ippc07 network connectivity was repaired. Added it to the ~ipp nebulous server list * Evidence is mounting that the problem with our throughput is in the nebulous database on ippdb00. Looking at the processlist it appears that queries are getting blockdd for several seconds. * 10:50 increased camera poll limit to 100. Since these processes are taking the better part of an hour to do the replication it doesn't hurt to get them all started. * 11:45 stopped processing in preparation for switching nebulous to the mysql on ippdb04 * 12:14 change of plans. starting processing again === Thursday : 2013.04.18 === === Friday : 2013.04.19 === === Saturday : 2013.04.20 === === Sunday : YYYY.MM.DD ===