| Version 15 (modified by , 14 years ago) ( diff ) |
|---|
(Up to PS1 IPP Czar Logs)
Monday : 2012-10-15
- 07:30 Bill: ipp020 is having nfs problems which has clogged stdscience and registration. Fixed the registration problem (we are a couple of hundred exposures behind). Working on ipp020 now.
- 07:50 Restarted stdscience and summitcopy. A stuck job in summitcopy was blocking burntool processing. We have 274 exposures copied but not registered. ipp020 is set to repair and set to off in registration and summit copy.
- 08:00 removed LAP label from stdscience to give nightlyscience all of the horespower
- 11:36 added lap label back. the database replication problem happened again and czartool is behind the times
- 14:14 Bill ran stop slave; set global sql_slave_skip_counter=1; start slave; on ippdb03 and ipp001 to fix the replication problem
- 22:10 MEH: looks like ipp016 has stalling registration and processing.. trying to isolate + clean up
- restarted registration w/o ipp016 - ok
- restarted stdscience w/o ipp016 - nightly_science.pl --queue_diffs hanging again.. will probably cause replication error
- restarted pstamp w/o ipp016 - ok
- neb-host ipp016 repair -- dont want to reboot tonight if don't have too, can still access /data/ipp016.0 just cannot ssh into (like ipp023, ipp013 etc recently)
Tuesday : 2012-10-16
- 06:08 EAM : ipp023 crashed, rebooted it (Ipp023-crash-20121016)
- 06:20 EAM : set ipp023 to 'retry' in summitcopy and registration (had been 'down' due to crash). also set ipp054 - ipp059 to 'on' in summitcopy. we used to have 30 machines doing the download, but now we seem to be down to 23 by default. I suspect this is because of concern about some of the wave 1 nodes. I think we should keep enough nodes in the list to keep the download rate acceptable.
Wednesday : 2012-10-17
Thursday : 2012-10-18
Friday : 2012-10-19
Saturday : 2012-10-20
Sunday : 2012-10-21
Note:
See TracWiki
for help on using the wiki.
