Server Admin Log
Appearance
2026-09-05
- 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-09-04
- 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
- 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
- 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
- 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
- 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
- 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
- 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
- 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
- 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
- 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
- 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
- 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
- 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
- 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
- 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
- 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
- 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
- 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
- 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
- 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
- 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
- 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
- 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
- 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
- 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
- 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
- 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
- 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
- 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
- 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
- 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
- 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
- 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
- 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
- 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
- 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
- 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
- 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
- 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
- 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
- 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
- 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
- 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
- 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
- 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
- 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
- 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
- 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
- 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
- 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
- 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
- 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
- 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
- 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
- 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
- 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
- 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
- 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
- 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
- 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
- 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
- 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
- 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
- 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
- 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
- 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
- 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
- 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
- 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
- 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
- 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
- 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
- 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
- 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
- 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
- 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
- 10:00 marostegui@cumin1003: Removing db1182 from zarcillo T434869
- 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
- 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
- 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
- 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
- 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
- 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
- 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
- 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl T434869', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
- 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
- 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
- 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
- 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
- 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
- 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
- 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
- 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
- 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
- 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
- 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
- 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
- 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
- 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
- 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
- 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
- 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
- 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
- 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
- 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
- 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
- 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for T436913 (duration: 34m 26s)
- 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
- 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
- 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
- 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
- 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
- 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
- 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
- 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
- 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
- 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
- 08:07 btullis@deploy1003: Started scap sync-world: Trying again for T436913
- 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
- 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
- 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
- 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
- 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
- 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
- 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
- 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
- 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
- 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
- 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
- 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
- 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
- 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
- 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
- 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
- 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
- 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
- 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
- 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
- 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
- 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
- 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
- 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
2026-09-03
- 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # T435363
- 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
- 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for Add exclusions to Apple app site association file (T435363) (duration: 12m 24s)
- 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
- 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
- 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
- 20:46 arlolra@deploy1003: arlolra, tsev: Backport for Add exclusions to Apple app site association file (T435363) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
- 20:42 arlolra@deploy1003: Started scap sync-world: Backport for Add exclusions to Apple app site association file (T435363)
- 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
- 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
- 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919) (duration: 10m 23s)
- 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
- 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
- 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
- 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
- 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
- 20:28 arlolra@deploy1003: Started scap sync-world: Backport for prv: Enable parsoid rendering for more wikisource wikis (T436919)
- 20:23 catrope@deploy1003: Finished scap sync-world: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135) (duration: 13m 41s)
- 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
- 20:16 catrope@deploy1003: catrope: Continuing with deployment
- 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
- 20:13 catrope@deploy1003: catrope: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
- 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
- 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
- 20:09 catrope@deploy1003: Started scap sync-world: Backport for Email confirmation A/A test: make registration cutoff consistent (T435135), Instrumentation for email confirmation upfront enforcement A/A test (T435135), Email confirmation A/A: check creation wiki, centralize logic (T435135)
- 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
- 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
- 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - T435419 (duration: 02m 59s)
- 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - T435419
- 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
- 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
- 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
- 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs T430837
- 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
- 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
- 17:39 ryankemper: T421642 [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
- 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
- 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
- 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
- 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
- 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
- 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
- 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
- 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
- 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
- 17:01 btullis@deploy1003: Started scap sync-world: Trying again for T436913
- 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
- 16:58 dancy: Running scap clean-images on deploy1003
- 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
- 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 16:39 btullis@deploy1003: Started scap sync-world: Trying again for T436913
- 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
- 16:14 ryankemper: T430880 Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
- 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
- 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
- 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
- 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
- 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for T436913
- 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for Growth: Remove now removed config variable (duration: 09m 29s)
- 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
- 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for Growth: Remove now removed config variable
- 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
- 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
- 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
- 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
- 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
- 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
- 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
- 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
- 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
- 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
- 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
- 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
- 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
- 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
- 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
- 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
- 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
- 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
- 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
- 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
- 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
- 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": T425441
- 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
- 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
- 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
- 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
- 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" T425441
- 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
- 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
- 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
- 14:19 arnaudb@dns1006: END - running authdns-update
- 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
- 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
- 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
- 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
- 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
- 14:17 arnaudb@dns1006: START - running authdns-update
- 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
- 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
- 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
- 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
- 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
- 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
- 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
- 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
- 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
- 14:07 samtar@deploy1003: Finished scap sync-world: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087) (duration: 09m 36s)
- 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
- 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
- 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
- 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
- 14:02 samtar@deploy1003: btullis, samtar: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
- 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
- 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
- 13:58 samtar@deploy1003: Started scap sync-world: Backport for Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)
- 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
- 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
- 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
- 13:51 moritzm: installing sqlite3 security updates
- 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
- 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
- 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
- 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
- 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
- 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
- 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
- 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
- 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
- 13:42 samtar@deploy1003: Finished scap sync-world: Backport for Remove revisionId from term fallback cache lines for properties (T434204) (duration: 13m 50s)
- 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
- 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
- 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
- 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
- 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
- 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
- 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
- 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
- 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
- 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
- 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for Remove revisionId from term fallback cache lines for properties (T434204) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
- 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
- 13:28 samtar@deploy1003: Started scap sync-world: Backport for Remove revisionId from term fallback cache lines for properties (T434204)
- 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
- 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
- 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
- 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
- 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:14 moritzm: installing bash updates from trixie point release
- 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
- 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
- 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
- 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
- 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
- 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
- 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
- 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
- 13:08 jelto@dns1004: END - running authdns-update
- 13:06 jelto@dns1004: START - running authdns-update
- 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
- 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
- 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
- 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
- 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
- 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
- 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
- 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
- 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
- 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
- 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
- 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
- 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
- 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
- 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
- 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
- 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
- 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
- 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
- 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
- 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
- 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
- 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
- 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
- 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
- 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
- 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
- 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
- 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
- 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
- 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
- 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
- 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
- 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
- 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
- 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
- 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
- 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
- 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
- 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
- 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
- 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
- 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
- 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
- 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
- 11:58 kart_: cxserver: Use urldownloader LVS endpoint (T429175)
- 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
- 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
- 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
- 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
- 11:55 moritzm: installing rsync security updates
- 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
- 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
- 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
- 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
- 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
- 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
- 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
- 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
- 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
- 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
- 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
- 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
- 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
- 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
- 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
- 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
- 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
- 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
- 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
- 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
- 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
- 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
- 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
- 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
- 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
- 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
- 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
- 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
- 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
- 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
- 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
- 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
- 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
- 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 T434778
- 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
- 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
- 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
- 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
- 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
- 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
- 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
- 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
- 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
- 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
- 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
- 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
- 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
- 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
- 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
- 09:39 hnowlan: fixed currently oncall pane in klaxon
- 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
- 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 T434751
- 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
- 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
- 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 T434775
- 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
- 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
- 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
- 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
- 09:30 zabe@deploy1003: Finished scap sync-world: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791) (duration: 09m 30s)
- 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 09:25 zabe@deploy1003: zabe: Continuing with deployment
- 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
- 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
- 09:25 zabe@deploy1003: zabe: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl T436904', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
- 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
- 09:21 zabe@deploy1003: Started scap sync-world: Backport for HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)
- 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
- 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
- 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable T435810
- 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
- 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
- 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
- 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 T434775
- 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
- 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
- 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
- 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
- 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
- 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
- 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
- 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
- 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 T434776
- 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
- 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
- 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server (duration: 01m 21s)
- 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to prod server
- 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server (duration: 01m 28s)
- 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): T436812 to backup server
- 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
- 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
- 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
- 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 T434287
- 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
- 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
- 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
- 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
- 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
- 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
- 07:29 chlod: UTC morning backport window done
- 07:27 chlod@deploy1003: Finished scap sync-world: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651) (duration: 11m 54s)
- 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
- 07:22 XioNoX: push pfw policies - T436729
- 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:15 chlod@deploy1003: Started scap sync-world: Backport for core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)
- 07:15 marostegui: Power off db1228 for maintenance
- 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
- 07:01 arnaudb@dns1006: END - running authdns-update
- 06:58 arnaudb@dns1006: START - running authdns-update
- 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
- 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
- 06:46 moritzm: installing libxml2 security updates
- 06:27 hashar: Upgrading CI Jenkins on contint1003 # T436812
- 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
- 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
- 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
- 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
- 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
- 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
- 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
- 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
2026-09-02
- 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
- 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for private/readme.php: Remove now removed secrets (T436880) (duration: 10m 21s)
- 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
- 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for private/readme.php: Remove now removed secrets (T436880) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for private/readme.php: Remove now removed secrets (T436880)
- 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
- 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
- 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
- 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
- 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
- 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
- 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for Campaigns should not override existing campaign query strings (T436681) (duration: 11m 03s)
- 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 22:30 jdlrobson@deploy1003: jdlrobson: Backport for Campaigns should not override existing campaign query strings (T436681) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for Campaigns should not override existing campaign query strings (T436681)
- 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054) (duration: 14m 14s)
- 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
- 21:55 krinkle@deploy1003: krinkle: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:50 krinkle@deploy1003: Started scap sync-world: Backport for Extract MathJax DOM filter into a separate file (T435274), Respect contextual binomial sizing (T434477 T418144 T401718), Remove the last vestiges of $wgVirtualRestConfig (T436054)
- 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes T431608
- 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637) (duration: 09m 48s)
- 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
- 21:34 jforrester@deploy1003: jforrester: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes T431608
- 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [WikiLambda] Log the …Orchestrator and …AbstractClient channels too, wikifunctions: Set up the functionmaintainer right for the community (T435637)
- 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
- 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
- 20:18 dancy@deploy1003: Started scap sync-world: testing
- 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
- 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
- 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851) (duration: 64m 27s)
- 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade T431608
- 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
- 18:57 jforrester@deploy1003: jforrester: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
- 18:53 jforrester@deploy1003: Started scap sync-world: Backport for Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)
- 18:31 sukhe@dns1004: END - running authdns-update
- 18:28 sukhe@dns1004: START - running authdns-update
- 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs T430837
- 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
- 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
- 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
- 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
- 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
- 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
- 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
- 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
- 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
- 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
- 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
- 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
- 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
- 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
- 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
- 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
- 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
- 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
- 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781 (duration: 48m 09s)
- 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - T436781
- 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
- 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
- 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
- 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
- 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
- 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
- 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
- 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
- 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
- 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
- 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
- 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
- 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
- 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
- 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
- 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
- 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
- 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
- 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
- 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
- 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
- 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
- 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
- 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
- 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
- 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
- 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
- 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
- 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
- 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
- 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
- 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781
- 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
- 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
- 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
- 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
- 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
- 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
- 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
- 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
- 14:32 moritzm: installing pdns-recursor security updates
- 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 T427897
- 13:36 samtar@deploy1003: Finished scap sync-world: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970) (duration: 09m 52s)
- 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
- 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:26 samtar@deploy1003: Started scap sync-world: Backport for Add thumb.* to allowed hosts for commons images (T436579), Add thumb.* to allowed hosts for commons images (T436579), Enable mobile MMV on all wikis (T429970)
- 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
- 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
- 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
- 13:24 moritzm: installing wireshark security updates
- 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
- 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
- 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
- 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
- 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
- 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia T436505
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
- 12:44 atsuko@dns1004: END - running authdns-update
- 12:41 atsuko@dns1004: START - running authdns-update
- 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
- 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
- 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
- 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517) (duration: 12m 50s)
- 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
- 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
- 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
- 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645), EventStreamConfig: Register the abuse_review_interaction stream (T435517)
- 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
- 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
- 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
- 11:31 marostegui@cumin1003: Removing db1172 from zarcillo T436763
- 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
- 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
- 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
- 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
- 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
- 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
- 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
- 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
- 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
- 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
- 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
- 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
- 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
- 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
- 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
- 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
- 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
- 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
- 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
- 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
- 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
- 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
- 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
- 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
- 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
- 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
- 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
- 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl T436763', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
- 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
- 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
- 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for T417800 (duration: 05m 37s)
- 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for T417800
- 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
- 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
- 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
- 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
- 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:56 jmm@dns1004: END - running authdns-update
- 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:53 jmm@dns1004: START - running authdns-update
- 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
- 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl T435892', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
- 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
- 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
- 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:30 moritzm: installing openjdk-21 security updates
- 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
- 09:10 moritzm: installing openjdk-8 security updates
- 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
- 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 08:46 tappof: bump space for prometheus k8s-dse in eqiad
- 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
- 08:36 moritzm: installing libgraphite2 security updates
- 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 08:23 Msz2001: UTC morning backport window done
- 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for Update stream config for user_info_card_interaction (T435585) (duration: 14m 36s)
- 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
- 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
- 08:14 mszwarc@deploy1003: mszwarc: Backport for Update stream config for user_info_card_interaction (T435585) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
- 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
- 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
- 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for Update stream config for user_info_card_interaction (T435585)
- 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
- 08:07 fabfur: depooling and silencing cp5022 (T414411)
- 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
- 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)}}
- 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
- 07:49 mszwarc@deploy1003: mszwarc: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585) synced to the
- 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
- 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
- 07:30 jmm@dns1004: END - running authdns-update
- {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the source page with the api_request event (T435585), UserInfoCard: Send the place of the trigger with api_request (T435585), [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
- 07:27 jmm@dns1004: START - running authdns-update
- 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734) (duration: 16m 04s)
- 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
- 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
- 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
- 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
- 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
- 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
- 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672), InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712), Add two wmf groups to privileged status (T436734)
- 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
- 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
- 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
- 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
- 06:23 slyngshede@dns1004: END - running authdns-update
- 06:21 marostegui: Drop cu* tables from s3 bswiktionary T435965
- 06:20 slyngshede@dns1004: START - running authdns-update
- 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
- 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - T436675
- 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses (duration: 04m 42s)
- 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
- 05:03 tstarling@deploy1003: tstarling: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 05:01 tstarling@deploy1003: Started scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses
- 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
- 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
- 04:29 tstarling@deploy1003: tstarling: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 04:25 tstarling@deploy1003: Started scap sync-world: Backport for Updater: Normalize MW_VERSION (T436741), Set a short CC:max-age on cacheable REST responses
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-09-01
- 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for Merge branch 'master' into wmf_deploy (duration: 18m 05s)
- 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 21:47 jdlrobson@deploy1003: jdlrobson: Backport for Merge branch 'master' into wmf_deploy synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for Merge branch 'master' into wmf_deploy
- 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248) (duration: 23m 55s)
- 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 21:20 jdlrobson@deploy1003: jdlrobson: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for Merge branch 'master' into wmf_deploy, Preserve showlogin query parameter on redirects (T435248)
- 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
- 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
- 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
- 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
- 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
- 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
- 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
- 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
- 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
- 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
- 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
- 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
- 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
- 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
- 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
- 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
- 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
- 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
- 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
- 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
- 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
- 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
- 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
- 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
- 19:06 sukhe@dns1004: END - running authdns-update
- 19:03 sukhe@dns1004: START - running authdns-update
- 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs T430837
- 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade T431608
- 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
- 16:52 dancy@deploy1003: Started scap sync-world: testing
- 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
- 16:48 dancy@deploy1003: Started scap sync-world: testing
- 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
- 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
- 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
- 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
- 15:51 moritzm: installing mesa security updates
- 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
- 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
- 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
- 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
- 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
- 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
- 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
- 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
- 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
- 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
- 14:55 hashar: Restarted Jenkins on releases1003
- 14:51 hashar: Restarted CI Jenkins on contint1003
- 14:48 hashar: Restarting Gerrit primary on gerrit2003
- 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
- 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
- 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
- 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
- 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
- 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
- 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
- 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
- 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
- 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
- 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
- 14:24 moritzm: installing curl security updates
- 14:24 jmm@dns1004: END - running authdns-update
- 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
- 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
- 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
- 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
- 14:21 jmm@dns1004: START - running authdns-update
- 14:21 jmm@dns1004: END - running authdns-update
- 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
- 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
- 14:19 jmm@dns1004: START - running authdns-update
- 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
- 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
- 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
- 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
- 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
- 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
- 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
- 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
- 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
- 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
- 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
- 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
- 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
- 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
- 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
- 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
- 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
- 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
- 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
- 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490) (duration: 37m 37s)
- 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
- 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
- 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # T418521
- 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
- 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
- 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
- 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
- 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
- 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
- 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
- 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
- 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
- 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
- 13:58 ladsgroup@dns1004: END - running authdns-update
- 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
- 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:56 ladsgroup@dns1004: START - running authdns-update
- 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 13:56 ladsgroup@dns1004: END - running authdns-update
- 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 13:53 ladsgroup@dns1004: START - running authdns-update
- 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
- 13:48 kharlan@deploy1003: kharlan: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
- 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 13:41 jmm@dns1004: END - running authdns-update
- 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
- 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
- 13:38 jmm@dns1004: START - running authdns-update
- 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
- 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
- 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
- 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
- 13:26 kharlan@deploy1003: Started scap sync-world: Backport for Special:AbuseReview: Show changes as core's inline diff (T436490)
- 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
- 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 13:23 aude@deploy1003: Finished scap sync-world: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922) (duration: 20m 24s)
- 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
- 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
- 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
- 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
- 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
- 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
- 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
- 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
- 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
- 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
- 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM T429175
- 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
- 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
- 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
- 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
- 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
- 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
- 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
- 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
- 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
- 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
- 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
- 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
- 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
- 13:11 aude@deploy1003: aude: Continuing with deployment
- 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
- 13:07 aude@deploy1003: aude: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
- 13:02 aude@deploy1003: Started scap sync-world: Backport for Enable ReadingLists for logged-in users on phase 1 wikis (T434922)
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
- 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
- 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
- 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
- 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
- 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
- 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
- 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
- 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
- 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848) (duration: 16m 25s)
- 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
- 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
- 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
- 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
- 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
- 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
- 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
- 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
- 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
- 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
- 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
- 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
- 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
- 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
- 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
- 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
- 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
- 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
- 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
- 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
- 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)
- 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
- 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
- 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
- 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
- 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
- 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
- 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
- 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
- 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
- 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
- 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
- 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
- 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
- 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
- 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
- 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
- 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
- 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
- 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
- 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
- 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
- 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
- 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
- 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
- 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
- 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
- 11:12 moritzm: installing openjdk-21 security updates
- 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
- 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
- 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
- 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
- 10:45 moritzm: installing Python 3.11 security updates
- 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
- 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
- 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
- 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
- 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
- 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
- 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
- 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
- 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
- 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
- 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
- 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
- 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
- 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
- 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
- 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
- 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
- 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
- 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
- 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
- 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
- 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
- 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
- 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
- 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production T427059', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
- 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
- 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
- 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
- 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
- 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
- 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
- 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
- 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
- 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
- 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
- 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
- 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
- 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
- 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
- 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 T427059', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
- 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
- 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
- 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
- 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
- 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
- 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
- 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
- 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
- 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
- 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
- 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
- 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
- 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
- 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
- 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
- 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
- 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
- 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
- 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
- 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
- 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
- 06:29 moritzm: installing Java 17 security updates
- 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
- 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
- 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
- 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs T430837 (duration: 37m 30s)
- 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs T430837
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
2026-08-31
- 23:50 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 23:50 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 23:49 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 23:49 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 23:13 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 23:12 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 22:49 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 22:49 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 21:30 sbassett: Deployed security fix for T435863
- 20:59 cdanis@deploy1003: helmfile [codfw] DONE helmfile.d/services/chart-renderer: apply
- 20:58 cdanis@deploy1003: helmfile [codfw] START helmfile.d/services/chart-renderer: apply
- 20:58 cdanis@deploy1003: helmfile [eqiad] DONE helmfile.d/services/chart-renderer: apply
- 20:58 cdanis@deploy1003: helmfile [eqiad] START helmfile.d/services/chart-renderer: apply
- 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # T432412
- 20:34 cdanis: kubectl -n chart-renderer scale deployment chart-renderer-production --replicas 4
- 20:33 arlolra@deploy1003: Finished scap sync-world: Backport for Add main page exclusions to Apple App Site Association File (T432412) (duration: 11m 04s)
- 20:28 arlolra@deploy1003: tsev, arlolra: Continuing with deployment
- 20:26 arlolra@deploy1003: tsev, arlolra: Backport for Add main page exclusions to Apple App Site Association File (T432412) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:22 arlolra@deploy1003: Started scap sync-world: Backport for Add main page exclusions to Apple App Site Association File (T432412)
- 20:18 arlolra@deploy1003: Finished scap sync-world: Backport for switch testwiki to use parsoid (T431636) (duration: 11m 29s)
- 20:11 arlolra@deploy1003: peterxy12, arlolra: Continuing with deployment
- 20:10 arlolra@deploy1003: peterxy12, arlolra: Backport for switch testwiki to use parsoid (T431636) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:06 arlolra@deploy1003: Started scap sync-world: Backport for switch testwiki to use parsoid (T431636)
- 18:20 ladsgroup@deploy1003: Finished scap sync-world: Backport for Enable thumb.wikimedia.org everywhere except enwiki (T427465) (duration: 11m 53s)
- 18:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 18:12 ladsgroup@deploy1003: ladsgroup: Backport for Enable thumb.wikimedia.org everywhere except enwiki (T427465) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:08 ladsgroup@deploy1003: Started scap sync-world: Backport for Enable thumb.wikimedia.org everywhere except enwiki (T427465)
- 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 17:53 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 17:43 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 17:42 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 17:41 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:41 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:32 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:32 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 17:32 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 17:32 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 16:29 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker[2006-2009].codfw.wmnet} and (A:aux-master-codfw or A:aux-worker-codfw)
- 16:29 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet
- 16:29 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet
- 16:24 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet
- 16:23 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet
- 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet
- 16:23 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet
- 16:18 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet
- 16:13 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet
- 16:12 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet
- 16:12 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet
- 16:07 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet
- 16:07 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet
- 16:07 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet
- 16:07 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet
- 16:06 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie
- 16:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet
- 16:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet
- 16:01 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker[2006-2009].codfw.wmnet} and (A:aux-master-codfw or A:aux-worker-codfw)
- 15:32 James_F: Deployed patch for T435086
- 15:28 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:sessionstore: Set storage compatability to NONE — T435154 - eevans@cumin1003
- 15:17 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:sessionstore: Set storage compatability to NONE — T435154 - eevans@cumin1003
- 15:10 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
- 15:09 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
- 15:09 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
- 15:09 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
- 15:09 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
- 15:08 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
- 14:48 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:sessionstore: Set storage compatability to UPGRADING — T435154 - eevans@cumin1003
- 14:45 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
- 14:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.* [reason: update IP addrs]
- 14:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:sessionstore: Set storage compatability to UPGRADING — T435154 - eevans@cumin1003
- 14:13 Emperor: attempt xfs_repair /dev/sdd1 on ms-be1090
- 13:21 moritzm: installing apr-util security updates
- 13:18 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 13:18 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 12:29 moritzm: installing openjdk-17 security updates
- 11:28 kamila@deploy1003: Finished scap sync-world: rebuild for T436488 (duration: 34m 43s)
- 11:21 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 11:20 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 11:11 moritzm: powercycle netmon2002, unresponsive
- 10:54 kamila@deploy1003: Started scap sync-world: rebuild for T436488
- 10:30 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker[1006-1009].eqiad.wmnet} and (A:aux-master-eqiad or A:aux-worker-eqiad)
- 10:30 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet
- 10:30 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet
- 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet
- 10:24 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet
- 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet
- 10:24 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet
- 10:19 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet
- 10:18 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet
- 10:18 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet
- 10:18 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet
- 10:15 ozge@deploy1003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply
- 10:14 ozge@deploy1003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply
- 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet
- 10:12 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet
- 10:12 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet
- 10:12 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet
- 10:12 ozge@deploy1003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply
- 10:10 ozge@deploy1003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply
- 10:07 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet
- 10:07 fabfur: enable puppet on A:cp to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330314 (T434766)
- 10:07 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet
- 10:07 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker[1006-1009].eqiad.wmnet} and (A:aux-master-eqiad or A:aux-worker-eqiad)
- 10:02 ozge@deploy1003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply
- 10:01 ozge@deploy1003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply
- 09:59 fabfur: disabled puppet on A:cp to gradually apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330314, starting from cp6016
- 09:47 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-master-codfw
- 09:47 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet
- 09:47 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet
- 09:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet
- 09:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet
- 09:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet
- 09:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet
- 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet
- 09:37 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet
- 09:37 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-master-codfw
- 09:34 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-master-eqiad
- 09:34 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet
- 09:34 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet
- 09:30 elukey@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=pki,name=eqiad
- 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet
- 09:29 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet
- 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet
- 09:29 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet
- 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host pki1002.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 09:24 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet
- 09:24 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet
- 09:24 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-master-eqiad
- 09:21 elukey@cumin1003: START - Cookbook sre.hosts.provision for host pki1002.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 09:16 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts pki1002.eqiad.wmnet
- 09:16 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki1002.eqiad.wmnet
- 09:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki1002.eqiad.wmnet
- 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 22 hosts with reason: Cloning
- 08:58 marostegui: Stop mariadb on sanitarium s2,s4,s6,s7 there will be lag on wikireplicas for those sections T407942
- 08:57 hashar: Upgrading CI Jenkins 2.555.3 to 2.568.2
- 08:56 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe
- 08:49 moritzm: installing emacs security updates
- 08:43 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe
- 08:42 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts pki1002.eqiad.wmnet
- 08:39 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe
- 08:39 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe
- 08:01 moritzm: installing Linux 6.12.107 on Trixie hosts
- 07:41 kartik@deploy1003: Finished scap sync-world: Backport for thwikibooks: update wordmark and tagline (T436426) (duration: 07m 41s)
- 07:38 kartik@deploy1003: hamishz, kartik: Rolling back deployment
- 07:37 kartik@deploy1003: hamishz, kartik: Backport for thwikibooks: update wordmark and tagline (T436426) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:33 kartik@deploy1003: Started scap sync-world: Backport for thwikibooks: update wordmark and tagline (T436426)
- 07:31 moritzm: installing openssl security updates
- 07:30 kartik@deploy1003: Finished scap sync-world: Backport for kowiki: convert extendedconfirmed calculation to begin from first edit (T436449) (duration: 15m 00s)
- 07:22 kartik@deploy1003: kartik, revi: Continuing with deployment
- 07:21 jelto@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
- 07:20 jelto@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
- 07:19 kartik@deploy1003: kartik, revi: Backport for kowiki: convert extendedconfirmed calculation to begin from first edit (T436449) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:15 kartik@deploy1003: Started scap sync-world: Backport for kowiki: convert extendedconfirmed calculation to begin from first edit (T436449)
- 07:05 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 07:04 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 03:44 tstarling@deploy1003: Finished scap sync-world: Backport for Produnto: use wgCopyUploadProxy to contact GitLab (T421436) (duration: 33m 20s)
- 03:31 tstarling@deploy1003: tstarling: Continuing with deployment
- 03:29 tstarling@deploy1003: tstarling: Backport for Produnto: use wgCopyUploadProxy to contact GitLab (T421436) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 03:11 tstarling@deploy1003: Started scap sync-world: Backport for Produnto: use wgCopyUploadProxy to contact GitLab (T421436)
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 15s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-08-30
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 34s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-08-29
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-08-28
- 21:16 ryankemper: T415073 admin_ng namespace-certificates applied on `staging-codfw`, `staging-eqiad`, `codfw`, `eqiad`; `wikidata-query-gui` cert reissued without `query-legacy-full.wikidata.org` SAN, verified on the wire in both DCs
- 21:08 ryankemper@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
- 21:07 ryankemper@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
- 21:07 ryankemper@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 21:07 ryankemper@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 21:07 ryankemper@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 21:07 ryankemper@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 21:06 ryankemper@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 21:06 ryankemper@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 21:05 ryankemper: T415073 beginning admin_ng namespace-certificates rollout (gerrit 1278562, drop query-legacy-full.wikidata.org SAN from wikidata-query-gui cert), starting with staging
- 20:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 20:04 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:47 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:45 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:45 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:44 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:43 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
- 19:31 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
- 18:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 18:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
- 18:09 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
- 17:53 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 17:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 17:50 eevans@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 17:50 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 17:32 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
- 16:38 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie
- 16:03 sukhe: sudo ipmitool -I lanplus -H "cp5022.mgmt.eqsin.wmnet" -U root -E chassis power cycle: trying to debug why it doesn't come up after reimage
- 15:43 elukey: powercycle pki1002 after CPU-related failures (host completely unresponsive) - T434268
- 15:40 elukey@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=pki,name=eqiad
- 15:03 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 15:03 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 14:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
- 14:54 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
- 14:22 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5022
- 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.move-vlan for host cp5022
- 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
- 14:03 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 14:01 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host cumin1004.eqiad.wmnet
- 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cumin1004.eqiad.wmnet with OS trixie
- 12:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cumin1004.eqiad.wmnet with reason: host reimage
- 12:07 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cumin1004.eqiad.wmnet with reason: host reimage
- 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host cumin1004.eqiad.wmnet with OS trixie
- 11:43 moritzm: add new LDAP group cn=airflow-experiment-platform-ops,ou=groups,dc=wikimedia,dc=org T416709
- 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM cumin1004.eqiad.wmnet - jmm@cumin2003"
- 11:43 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM cumin1004.eqiad.wmnet - jmm@cumin2003"
- 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cumin1004.eqiad.wmnet on all recursors
- 11:41 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache cumin1004.eqiad.wmnet on all recursors
- 11:40 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:39 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM cumin1004.eqiad.wmnet - jmm@cumin2003"
- 11:38 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM cumin1004.eqiad.wmnet - jmm@cumin2003"
- 11:23 jmm@cumin2003: START - Cookbook sre.dns.netbox
- 11:23 jmm@cumin2003: START - Cookbook sre.ganeti.makevm for new host cumin1004.eqiad.wmnet
- 11:12 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T436069
- 10:43 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T436069
- 08:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin2002.codfw.wmnet
- 08:59 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 08:59 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
- 08:58 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
- 08:28 jmm@cumin2003: START - Cookbook sre.dns.netbox
- 07:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin2002.codfw.wmnet
- 06:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org
- 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org
- 04:34 ryankemper: T435862 Completed the rolling restart of all 15 production Presto workers; verified `true/true` spill settings across coordinators and workers, full runtime membership, the formerly failing query, and Superset dashboard 757. This is done
- 04:31 ryankemper@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
- 03:59 ryankemper@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
- 03:58 ryankemper: T435862 Restarted both production Presto coordinators (standby first, then active) after enabling JOIN spilling; the active coordinator reports `spill_enabled=true` and `join_spill_enabled=true` and all 15 workers rejoined
- 03:42 ryankemper: T435862 Verified `spill_enabled=true` and `join_spill_enabled=true` on the Presto test cluster after restarting its coordinator and worker; both nodes are active and distributed query `20260828_033922_00002_u9k78` succeeded
- 03:40 ryankemper@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
- 03:38 ryankemper@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
- 03:10 tstarling@deploy1003: Finished scap sync-world: Backport for Grant produnto-update to groups that have editprotected (T421436), Produnto: Add IPv6 range for GitLab (T421436) (duration: 12m 44s)
- 03:05 tstarling@deploy1003: tstarling: Continuing with deployment
- 03:02 tstarling@deploy1003: tstarling: Backport for Grant produnto-update to groups that have editprotected (T421436), Produnto: Add IPv6 range for GitLab (T421436) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 02:58 tstarling@deploy1003: Started scap sync-world: Backport for Grant produnto-update to groups that have editprotected (T421436), Produnto: Add IPv6 range for GitLab (T421436)
- 02:25 ryankemper: `ryankemper@pcc-db1002:~$ sudo -u jenkins-deploy puppetdb-populate --host an-test-coord1002.eqiad.wmnet` (refreshed stale catalog)
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 53s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:55 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 00:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 00:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 00:53 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 00:53 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 00:52 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 00:40 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 00:39 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 00:27 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 00:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 00:24 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 00:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
2026-08-27
- 23:14 jasmine@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1261.eqiad.wmnet
- 23:14 jasmine@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1261.eqiad.wmnet
- 23:14 jasmine@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1261.eqiad.wmnet
- 23:07 jasmine_: homer lsw1-c6-eqiad* commit 'T421711'
- 23:05 jasmine@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1261.eqiad.wmnet with OS trixie
- 22:44 jasmine@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1261.eqiad.wmnet with reason: host reimage
- 22:40 jasmine@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1261.eqiad.wmnet with reason: host reimage
- 22:20 jasmine@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1261
- 22:20 jasmine@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1261
- 22:19 jasmine@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1261
- 22:19 jasmine@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1261.eqiad.wmnet 71.32.64.10.in-addr.arpa 1.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 22:19 jasmine@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1261.eqiad.wmnet 71.32.64.10.in-addr.arpa 1.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 22:19 jasmine@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 22:19 jasmine@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1261 - jasmine@cumin1003"
- 22:19 jasmine@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1261 - jasmine@cumin1003"
- 22:13 jasmine@cumin1003: START - Cookbook sre.dns.netbox
- 22:12 jasmine@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1261
- 22:12 jasmine@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1261.eqiad.wmnet with OS trixie
- 22:11 jasmine@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1261.eqiad.wmnet
- 22:11 jasmine@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1261.eqiad.wmnet
- 22:11 jasmine@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1261.eqiad.wmnet
- 21:40 sbassett: Deployed security mitigations for T432713, T435026
- 20:47 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 20:47 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 20:37 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
- 20:31 krinkle@deploy1003: Finished scap sync-world: Backport for Fix incorrect number of children in m(under|over) (T435705) (duration: 17m 00s)
- 20:29 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
- 20:25 krinkle@deploy1003: krinkle: Continuing with deployment
- 20:22 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
- 20:18 krinkle@deploy1003: krinkle: Backport for Fix incorrect number of children in m(under|over) (T435705) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:14 krinkle@deploy1003: Started scap sync-world: Backport for Fix incorrect number of children in m(under|over) (T435705)
- 19:41 inflatador: [bking@ganeti1046] ~$ sudo gnt-instance replace-disks -n ganeti1057.eqiad.wmnet krb1004.eqiad.wmnet T435873
- 19:39 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 00s)
- 19:36 dancy@deploy1003: Started scap sync-world: testing
- 19:35 dancy@deploy1003: sync-world failed: <RuntimeError> dictionary changed size during iteration (scap version: 4.286.0) (duration: 01m 12s)
- 19:34 dancy@deploy1003: Started scap sync-world: testing
- 19:34 inflatador: [bking@ganeti1046] ~$ sudo gnt-instance migrate krb1004 ganeti1035 -> ganeti1029 T435873
- 19:28 dancy@deploy1003: Finished scap sync-world: testing (duration: 05m 45s)
- 19:23 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5022
- 19:23 cdobbins@cumin1003: START - Cookbook sre.hosts.move-vlan for host cp5022
- 19:23 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
- 19:22 dancy@deploy1003: Started scap sync-world: testing
- 19:22 dancy@deploy1003: Installation of scap version "4.286.0" completed for 3 hosts
- 19:20 dancy@deploy1003: Installing scap version "4.286.0" for 3 host(s)
- 18:37 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
- 18:36 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
- 18:36 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
- 18:36 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
- 18:36 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
- 18:35 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
- 18:34 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/termbox: apply
- 18:34 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/termbox: apply
- 18:33 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
- 18:25 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply