== PS1 IPP Czar Logs for the week 2011.06.27 - 2011.07.03 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday: 2011-06-27 === * 07:17 CZW: marked exposure impacted by loss of summit computer as drop to allow registration to proceed: {{{pztool -updatepzexp -exp_name o5739g0084o -inst gpc1 -telescope ps1 -set_state drop -summit_id 348983}}} * 08:00 (Roy) 3 exposures not registered. Looks like one or all of summitcopy/burntool/registration were down most of the night. * 10:58 serge: in stdscience: {{{ del.label LAP.ThreePi.20110621 }}} * 15:17 serge stopped gpc1 pantasks servers. heather stopped addstar and isp related ones * 15:53 serge started dump of gpc1 czardb isp ippadmin ipptopsps to /export/ippdb01.0/mysql_gpc1.backup/gpc1.20110627.sql.bz2 {{{ mysqldump [hidden connection information] --databases gpc1 czardb isp ippadmin | ~ipp/local/bin/pbzip2 > gpc1.20110627.sql.bz2 }}} About the master: {{{ mysql> SHOW MASTER STATUS; +-------------------+----------+--------------+------------------+ | File | Position | Binlog_Do_DB | Binlog_Ignore_DB | +-------------------+----------+--------------+------------------+ | mysqld-bin.020077 | 84637877 | | | +-------------------+----------+--------------+------------------+ }}} * 17:09 Patnasks servers have been restarted * 17:45 CZW: Finally remembered to queue the cleanup of LAP/20110505 distribution bundles for bad stacks. * 18:26 CZW: Rebuilt psModules to pick up stacking speed improvements. This may not be sufficient, however, so more work is likely needed. === Tuesday : 2011-06-28 === * 05:30 Serge: gpc1 is replicated on ippdb03 (ingestion on ipp001 is not finished yet) * 08:13 Roy: all exposures downloaded. Will add {{{LAP.ThreePi.20110621}}} label back-in when 3PI is done * 09:15 Serge: gpc1 is replicated on ipp001 * 09:35 Serge: gpc1, ippadmin, and isp dumps are distributed as they should be. I will keep an eye on them though * 12:00 roy: ipp033 is down on Ganglia, but has network. Gavin is on the case. * 13:50 roy: ipp033 is back, but /data/ipp033.0/ is not visible from other machines * 14:00 Serge: replication broken on both hosts. ippdb01 is caching the hostname of ippdb03 with the old hostname value (ippc00). I don't think it's the explanation for the failure though. I'm ingesting yesterday's dump in ippdb03 once again. I will not ingest it in ipp001. I leave the old ingestion schema activated (ingestion in gpc1_0 and gpc1_1). * 15:15 - 15:40 Bill: regenerated several corrupt files using the scripts in ipp/tools {{{ corrupt file on ipp031 caused diff failure --warp_id 215094 --skycell_id skycell.1276.049 corrupt file on ipp018 caused diff failure --warp_id 215061 --skycell_id skycell.0962.042 corrupt file on ipp051 caused warp failure --cam_id 226915 corrupt file on ipp008 caused magic failure --diff_id 141410 --skycell_id skycell.1992.136 corrupt file on ipp012 caused magic failure --diff_id 141396 --skycell_id skycell.1992.143 corrupt file on ipp012 caused magic failure --diff_id 141386 --skycell_id skycell.1992.134 corrupt file on ipp012 caused distribution failure. Fixing this required faulting the corrupted magicDSFile and setting the states in the DB back to new. }}} * 15:35 turned magic revert and dist revert off to track down some corrupt files === Wednesday : YYYY.MM.DD === === Thursday : YYYY.MM.DD === === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===