Server Admin Log/Archive 107
Appearance
2026-06-30
- 23:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2102.codfw.wmnet with reason: host reimage
- 23:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2103.codfw.wmnet with reason: host reimage
- 23:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2102.codfw.wmnet with reason: host reimage
- 23:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2103.codfw.wmnet with reason: host reimage
- 23:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2102.codfw.wmnet with OS trixie
- 23:30 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2103.codfw.wmnet with OS trixie
- 23:12 ryankemper: T429844 [opensearch] cleared stale/redundant cluster.routing.allocation.* transient settings from all production search clusters; also cleared redundant action.auto_create_index and cluster.routing.use_adaptive_replica_selection transients from codfw chi/psi, and stale transient DEBUG logger overrides from codfw chi
- 23:06 dr0ptp4kt@deploy1003: Finished deploy [analytics/refinery@4e7a2b3] (thin): Regular analytics weekly train THIN [analytics/refinery@4e7a2b32] (duration: 02m 00s)
- 23:04 dr0ptp4kt@deploy1003: Started deploy [analytics/refinery@4e7a2b3] (thin): Regular analytics weekly train THIN [analytics/refinery@4e7a2b32]
- 23:04 dr0ptp4kt@deploy1003: Finished deploy [analytics/refinery@4e7a2b3]: Regular analytics weekly train [analytics/refinery@4e7a2b32] (duration: 04m 11s)
- 23:00 dr0ptp4kt@deploy1003: Started deploy [analytics/refinery@4e7a2b3]: Regular analytics weekly train [analytics/refinery@4e7a2b32]
- 22:56 dr0ptp4kt@deploy1003: Finished deploy [analytics/refinery@4e7a2b3] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@4e7a2b32] (duration: 01m 59s)
- 22:54 dr0ptp4kt@deploy1003: Started deploy [analytics/refinery@4e7a2b3] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@4e7a2b32]
- 22:53 dr0ptp4kt: Deploying Refinery at 4e7a2b32 for changes: pageview allowlist 1305158 (+min.wikiquote) 1305162 (+bol.wikipedia), 1305156 (+isv.wikipedia); 1305980 (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop 1295064 (+globalimagelinks) 1295069 (+filerevision)
- 22:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2071.codfw.wmnet with OS trixie
- 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2071.codfw.wmnet with reason: host reimage
- 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2071.codfw.wmnet with reason: host reimage
- 21:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 21:35 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 21:35 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 21:35 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 21:27 cmooney@dns3003: END - running authdns-update
- 21:25 cmooney@dns3003: START - running authdns-update
- 21:20 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2071.codfw.wmnet with OS trixie
- 21:19 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 21:19 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to drmrs - cmooney@cumin1003"
- 21:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2113.codfw.wmnet with OS trixie
- 21:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to drmrs - cmooney@cumin1003"
- 21:15 tgr_: UTC late deploysdone
- 21:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2101.codfw.wmnet with OS trixie
- 21:14 Dreamy_Jazz: Destroyed NodeJS iPoid service deployments for T416623
- 21:13 tgr@deploy1003: Finished scap sync-world: Backport for SecurityLogs: Create by moving code from mediawiki-config (T430564), SecurityLogs: Add tests (T430564), Remove security-related log hooks (T430564) (duration: 17m 45s)
- 21:11 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 21:08 tgr@deploy1003: tgr: Continuing with deployment
- 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2113.codfw.wmnet with reason: host reimage
- 20:57 tgr@deploy1003: tgr: Backport for SecurityLogs: Create by moving code from mediawiki-config (T430564), SecurityLogs: Add tests (T430564), Remove security-related log hooks (T430564) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:55 tgr@deploy1003: Started scap sync-world: Backport for SecurityLogs: Create by moving code from mediawiki-config (T430564), SecurityLogs: Add tests (T430564), Remove security-related log hooks (T430564)
- 20:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2101.codfw.wmnet with reason: host reimage
- 20:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2113.codfw.wmnet with reason: host reimage
- 20:46 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2101.codfw.wmnet with reason: host reimage
- 20:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:41 cscott@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid Read views for 50% of English Wikipedia desktop traffic (T430194) (duration: 10m 24s)
- 20:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:36 cscott@deploy1003: cscott: Continuing with deployment
- 20:36 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2089.codfw.wmnet with OS trixie
- 20:35 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 20:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 20:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2113.codfw.wmnet with OS trixie
- 20:32 cscott@deploy1003: cscott: Backport for Turn on Parsoid Read views for 50% of English Wikipedia desktop traffic (T430194) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:30 cscott@deploy1003: Started scap sync-world: Backport for Turn on Parsoid Read views for 50% of English Wikipedia desktop traffic (T430194)
- 20:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2101.codfw.wmnet with OS trixie
- 20:25 ebernhardson@deploy1003: Finished scap sync-world: Backport for Revert^3 "cirrus: AB test query suggester variants" (T407432), Revert "nlwiki: change to Wikipedia 25 logo" (T424519), ExtensionDistributor: mark 1.46 as stable (T423272) (duration: 13m 43s)
- 20:18 ebernhardson@deploy1003: chlod, ebernhardson, ariel: Continuing with deployment
- 20:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2101.codfw.wmnet with OS trixie
- 20:17 ebernhardson@deploy1003: chlod, ebernhardson, ariel: Backport for Revert^3 "cirrus: AB test query suggester variants" (T407432), Revert "nlwiki: change to Wikipedia 25 logo" (T424519), ExtensionDistributor: mark 1.46 as stable (T423272) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:15 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2089.codfw.wmnet with reason: host reimage
- 20:11 ebernhardson@deploy1003: Started scap sync-world: Backport for Revert^3 "cirrus: AB test query suggester variants" (T407432), Revert "nlwiki: change to Wikipedia 25 logo" (T424519), ExtensionDistributor: mark 1.46 as stable (T423272)
- 20:11 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2089.codfw.wmnet with reason: host reimage
- 20:08 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:08 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2066.codfw.wmnet with OS trixie
- 19:56 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2082.codfw.wmnet with OS trixie
- 19:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2101.codfw.wmnet with reason: host reimage
- 19:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2101.codfw.wmnet with reason: host reimage
- 19:50 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2089.codfw.wmnet with OS trixie
- 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2066.codfw.wmnet with reason: host reimage
- 19:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2082.codfw.wmnet with reason: host reimage
- 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2066.codfw.wmnet with reason: host reimage
- 19:31 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2101.codfw.wmnet with OS trixie
- 19:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2082.codfw.wmnet with reason: host reimage
- 19:24 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cirrussearch2089.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2066.codfw.wmnet with OS trixie
- 19:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2082.codfw.wmnet with OS trixie
- 19:13 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host cirrussearch2089.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
- 18:55 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3081.*
- 18:54 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3081.esams.wmnet with OS trixie
- 18:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2065.codfw.wmnet with OS trixie
- 18:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2110.codfw.wmnet with OS trixie
- 18:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2079.codfw.wmnet with OS trixie
- 18:29 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3081.esams.wmnet with reason: host reimage
- 18:23 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3081.esams.wmnet with reason: host reimage
- 18:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2065.codfw.wmnet with reason: host reimage
- 18:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2110.codfw.wmnet with reason: host reimage
- 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2065.codfw.wmnet with reason: host reimage
- 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2079.codfw.wmnet with reason: host reimage
- 18:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2110.codfw.wmnet with reason: host reimage
- 18:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2079.codfw.wmnet with reason: host reimage
- 18:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 18:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Migration of db1263.eqiad.wmnet completed
- 17:58 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3081.esams.wmnet with OS trixie
- 17:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2065.codfw.wmnet with OS trixie
- 17:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2079.codfw.wmnet with OS trixie
- 17:48 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3081.*
- 17:47 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2110.codfw.wmnet with OS trixie
- 17:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors
- 17:44 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors
- 17:43 sukhe@dns1004: END - running authdns-update
- 17:41 sukhe@dns1004: START - running authdns-update
- 17:18 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Migration of db1263.eqiad.wmnet completed
- 16:52 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1263.eqiad.wmnet with OS trixie
- 16:44 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
- 16:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 16:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 16:43 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
- 16:40 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cirrussearch2089.codfw.wmnet']
- 16:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 16:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 16:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1263.eqiad.wmnet with reason: host reimage
- 16:34 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cirrussearch2089.codfw.wmnet']
- 16:32 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cirrussearch2073.codfw.wmnet with OS trixie
- 16:28 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1263.eqiad.wmnet with reason: host reimage
- 16:16 kharlan@deploy1003: Finished scap sync-world: Backport for ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892), ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892) (duration: 33m 43s)
- 16:11 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1263.eqiad.wmnet with OS trixie
- 16:02 kharlan@deploy1003: kharlan: Continuing with deployment
- 16:02 kharlan@deploy1003: kharlan: Backport for ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892), ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:42 kharlan@deploy1003: Started scap sync-world: Backport for ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892), ConfirmEdit: Show a clear message when a login CAPTCHA is missing (T428892)
- 15:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1263: Upgrading db1263.eqiad.wmnet
- 15:38 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1263: Upgrading db1263.eqiad.wmnet
- 15:37 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 15:37 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:36 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2097.codfw.wmnet with OS trixie
- 15:31 urbanecm@deploy1003: Finished scap sync-world: Backport for Revert "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386) (duration: 08m 01s)
- 15:29 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2096.codfw.wmnet with OS trixie
- 15:28 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2095.codfw.wmnet with OS trixie
- 15:27 urbanecm@deploy1003: urbanecm: Continuing with deployment
- 15:25 urbanecm@deploy1003: urbanecm: Backport for Revert "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:23 urbanecm@deploy1003: Started scap sync-world: Backport for Revert "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)
- 15:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1262: Migration of db1262.eqiad.wmnet completed
- 15:17 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # T427386
- 15:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2097.codfw.wmnet with reason: host reimage
- 15:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2096.codfw.wmnet with reason: host reimage
- 15:05 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2095.codfw.wmnet with reason: host reimage
- 15:00 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
- 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 30818
- 15:00 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
- 15:00 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2097.codfw.wmnet with reason: host reimage
- 15:00 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
- 14:59 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 30818
- 14:59 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
- 14:59 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync
- 14:59 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync
- 14:59 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2096.codfw.wmnet with reason: host reimage
- 14:58 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync
- 14:58 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync
- 14:58 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync
- 14:57 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync
- 14:57 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: sync
- 14:57 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: sync
- 14:57 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2095.codfw.wmnet with reason: host reimage
- 14:57 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
- 14:56 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
- 14:55 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
- 14:55 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
- 14:55 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
- 14:54 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
- 14:54 ottomata: roll restart eventgates to re-cache page_change related schemas - T423583#12067516
- 14:50 jforrester@deploy1003: Finished scap sync-world: Backport for abstractwiki: Show Abstract provenance notice to all readers, not just sysops (T422710), abstractwiki: Stop mis-setting the relevant title on Abstract surfaces (T422655), SkinComponentFooter: Show copyright for known, not just existing, pages (T422655) (duration: 06m 55s)
- 14:49 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
- 14:48 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
- 14:47 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-staging-worker@codfw
- 14:47 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 14:46 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 14:46 jforrester@deploy1003: jforrester: Continuing with deployment
- 14:45 jforrester@deploy1003: jforrester: Backport for abstractwiki: Show Abstract provenance notice to all readers, not just sysops (T422710), abstractwiki: Stop mis-setting the relevant title on Abstract surfaces (T422655), SkinComponentFooter: Show copyright for known, not just existing, pages (T422655) synced to the testservers (see https://wikitech.wikimedia.org/w
- 14:43 jforrester@deploy1003: Started scap sync-world: Backport for abstractwiki: Show Abstract provenance notice to all readers, not just sysops (T422710), abstractwiki: Stop mis-setting the relevant title on Abstract surfaces (T422655), SkinComponentFooter: Show copyright for known, not just existing, pages (T422655)
- 14:42 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-staging-worker@codfw
- 14:42 jforrester@deploy1003: Started scap sync-world: Backport for abstractwiki: Show Abstract provenance notice to all readers, not just sysops (T422710), abstractwiki: Stop mis-setting the relevant title on Abstract surfaces (T422655), SkinComponentFooter: Show copyright for known, not just existing, pages (T422655)
- 14:40 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2097.codfw.wmnet with OS trixie
- 14:38 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2096.codfw.wmnet with OS trixie
- 14:36 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2095.codfw.wmnet with OS trixie
- 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1262: Migration of db1262.eqiad.wmnet completed
- 14:27 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-staging-worker@codfw
- 14:27 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 14:26 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 14:23 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-staging-worker@codfw
- 14:18 Lucas_WMDE: UTC afternoon backport+config window done
- 14:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [Growth] frwiki: Deploy automated mentor list cleaner (T427386) (duration: 06m 50s)
- 14:10 urbanecm@deploy1003: urbanecm: Continuing with deployment
- 14:10 urbanecm@deploy1003: urbanecm: Backport for [Growth] frwiki: Deploy automated mentor list cleaner (T427386) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:08 urbanecm@deploy1003: Started scap sync-world: Backport for [Growth] frwiki: Deploy automated mentor list cleaner (T427386)
- 14:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1262.eqiad.wmnet with OS trixie
- 14:01 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2094.codfw.wmnet with OS trixie
- 14:00 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
- 13:56 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2070.codfw.wmnet with OS trixie
- 13:53 bpirkle@deploy1003: Finished scap sync-world: Backport for REST: remove obsolete and unnecessary config entries (T422770 T423058 T422771) (duration: 10m 21s)
- 13:52 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2078.codfw.wmnet with OS trixie
- 13:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1262.eqiad.wmnet with reason: host reimage
- 13:48 bpirkle@deploy1003: bpirkle: Continuing with deployment
- 13:45 bpirkle@deploy1003: bpirkle: Backport for REST: remove obsolete and unnecessary config entries (T422770 T423058 T422771) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:42 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1262.eqiad.wmnet with reason: host reimage
- 13:42 bpirkle@deploy1003: Started scap sync-world: Backport for REST: remove obsolete and unnecessary config entries (T422770 T423058 T422771)
- 13:39 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2094.codfw.wmnet with reason: host reimage
- 13:38 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for postEdit: temp account experiment instrumentation (T429110), maybeSendThankYouEdit: avoid sending notification to temp users (T429110 T424205) (duration: 15m 12s)
- 13:35 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2070.codfw.wmnet with reason: host reimage
- 13:34 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, migr: Continuing with deployment
- 13:30 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2078.codfw.wmnet with reason: host reimage
- 13:27 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2070.codfw.wmnet with reason: host reimage
- 13:27 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2094.codfw.wmnet with reason: host reimage
- 13:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1262.eqiad.wmnet with OS trixie
- 13:25 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, migr: Backport for postEdit: temp account experiment instrumentation (T429110), maybeSendThankYouEdit: avoid sending notification to temp users (T429110 T424205) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:23 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for postEdit: temp account experiment instrumentation (T429110), maybeSendThankYouEdit: avoid sending notification to temp users (T429110 T424205)
- 13:22 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2078.codfw.wmnet with reason: host reimage
- 13:11 ihurbain@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid Read views for 25% of English Wikipedia desktop traffic (T430194) (duration: 07m 29s)
- 13:09 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2070.codfw.wmnet with OS trixie
- 13:07 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2094.codfw.wmnet with OS trixie
- 13:06 ihurbain@deploy1003: ihurbain: Continuing with deployment
- 13:05 ihurbain@deploy1003: ihurbain: Backport for Turn on Parsoid Read views for 25% of English Wikipedia desktop traffic (T430194) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:04 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2078.codfw.wmnet with OS trixie
- 13:03 ihurbain@deploy1003: Started scap sync-world: Backport for Turn on Parsoid Read views for 25% of English Wikipedia desktop traffic (T430194)
- 13:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 12:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 12:57 klausman@cumin2002: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ml-cache[2001-2003].codfw.wmnet,ml-cache[1001-1003].eqiad.wmnet
- 12:56 klausman@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:56 klausman@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ml-cache[2001-2003].codfw.wmnet,ml-cache[1001-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - klausman@cumin2002"
- 12:56 klausman@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ml-cache[2001-2003].codfw.wmnet,ml-cache[1001-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - klausman@cumin2002"
- 12:51 klausman@cumin2002: START - Cookbook sre.dns.netbox
- 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2105.codfw.wmnet with reason: host reimage
- 12:10 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2104.codfw.wmnet with reason: host reimage
- 12:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1261: Migration of db1261.eqiad.wmnet completed
- 12:07 elukey: upgrade all bookworm hosts to pywmflib 3.0 - T430552
- 12:04 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2105.codfw.wmnet with reason: host reimage
- 12:02 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2104.codfw.wmnet with reason: host reimage
- 11:48 kharlan@deploy1003: Finished scap sync-world: Backport for SimpleCaptcha: Log skipcaptcha right in force-show trigger (T402595) (duration: 07m 39s)
- 11:44 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2105.codfw.wmnet with OS trixie
- 11:44 kharlan@deploy1003: kharlan: Continuing with deployment
- 11:43 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2104.codfw.wmnet with OS trixie
- 11:42 kharlan@deploy1003: kharlan: Backport for SimpleCaptcha: Log skipcaptcha right in force-show trigger (T402595) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:42 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1014.eqiad.wmnet,service=s7
- 11:42 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1014.eqiad.wmnet,service=s2
- 11:40 kharlan@deploy1003: Started scap sync-world: Backport for SimpleCaptcha: Log skipcaptcha right in force-show trigger (T402595)
- 11:39 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
- 11:38 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
- 11:36 mszwarc@deploy1003: Finished scap sync-world: Backport for SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Link to Interaction Timeline for shared pages (T429785) (duration: 07m 48s)
- 11:31 mszwarc@deploy1003: mszwarc: Continuing with deployment
- 11:30 mszwarc@deploy1003: mszwarc: Backport for SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Link to Interaction Timeline for shared pages (T429785) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebu
- 11:28 mszwarc@deploy1003: Started scap sync-world: Backport for SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Defer signal matching until transaction commits (T430617), SuggestedInvestigations: Link to Interaction Timeline for shared pages (T429785)
- 11:26 moritzm: installing libpng1.6 security updates
- 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1261: Migration of db1261.eqiad.wmnet completed
- 11:21 Tran: Deployed patch for T427287
- 11:13 moritzm: installing Linux 6.12.94 on Trixie hosts
- 11:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 11:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
- 10:56 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 10:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 10:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 10:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on dbproxy[1026,1028].eqiad.wmnet with reason: cloning
- 10:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: cloning
- 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 10:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1217.eqiad.wmnet with reason: cloning
- 10:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1261.eqiad.wmnet with OS trixie
- 10:33 javiermonton@deploy1003: Finished scap sync-world: Backport for stream: webrequest.page_view.dev0 (T426091) (duration: 07m 56s)
- 10:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply}
- 10:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
- 10:28 javiermonton@deploy1003: javiermonton: Continuing with deployment
- 10:27 javiermonton@deploy1003: javiermonton: Backport for stream: webrequest.page_view.dev0 (T426091) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:25 daniel@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 10:25 daniel@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 10:25 javiermonton@deploy1003: Started scap sync-world: Backport for stream: webrequest.page_view.dev0 (T426091)
- 10:20 daniel@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 10:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1261.eqiad.wmnet with reason: host reimage
- 10:18 daniel@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 10:15 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1261.eqiad.wmnet with reason: host reimage
- 10:12 daniel@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 10:07 daniel@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1261.eqiad.wmnet with OS trixie
- 09:56 godog: restart pybal on A:lvs-high-traffic2-eqiad
- 09:53 godog: restart pybal on A:lvs-secondary-eqiad
- 09:51 aikochou@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: sync
- 09:51 aikochou@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: sync
- 09:49 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2204 (T426633)', diff saved to https://phabricator.wikimedia.org/P94616 and previous config saved to /var/cache/conftool/dbconfig/20260630-094938-fceratto.json
- 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2204', diff saved to https://phabricator.wikimedia.org/P94615 and previous config saved to /var/cache/conftool/dbconfig/20260630-093931-fceratto.json
- 09:38 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs T423918
- 09:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1261: Upgrading db1261.eqiad.wmnet
- 09:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1261: Upgrading db1261.eqiad.wmnet
- 09:34 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 09:34 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 09:29 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2204', diff saved to https://phabricator.wikimedia.org/P94613 and previous config saved to /var/cache/conftool/dbconfig/20260630-092923-fceratto.json
- 09:24 aklapper@deploy1003: Finished scap sync-world: Backport for Fix overflow menu for non-advanced users (T428220) (duration: 11m 40s)
- 09:19 filippo@puppetserver1001: conftool action : set/pooled=yes:weight=100; selector: service=dumps-nfs
- 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2204 (T426633)', diff saved to https://phabricator.wikimedia.org/P94612 and previous config saved to /var/cache/conftool/dbconfig/20260630-091915-fceratto.json
- 09:17 aklapper@deploy1003: aklapper: Continuing with deployment
- 09:16 aklapper@deploy1003: aklapper: Backport for Fix overflow menu for non-advanced users (T428220) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:13 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2204 (T426633)', diff saved to https://phabricator.wikimedia.org/P94611 and previous config saved to /var/cache/conftool/dbconfig/20260630-091307-fceratto.json
- 09:13 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance
- 09:12 aklapper@deploy1003: Started scap sync-world: Backport for Fix overflow menu for non-advanced users (T428220)
- 09:08 fceratto@cumin1003: dbctl commit (dc=all): 'Set weight db2204 T430624', diff saved to https://phabricator.wikimedia.org/P94610 and previous config saved to /var/cache/conftool/dbconfig/20260630-090841-fceratto.json
- 09:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary T430624', diff saved to https://phabricator.wikimedia.org/P94609 and previous config saved to /var/cache/conftool/dbconfig/20260630-090530-fceratto.json
- 09:04 federico3: Starting s2 codfw failover from db2204 to db2207 - T430624
- 09:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2
- 09:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7
- 09:02 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1014,1027].eqiad.wmnet with reason: cloning
- 09:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1027.eqiad.wmnet,service=s7
- 09:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1027.eqiad.wmnet,service=s2
- 08:57 jmm@dns1004: END - running authdns-update
- 08:55 jmm@dns1004: START - running authdns-update
- 08:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1223 (T426633)', diff saved to https://phabricator.wikimedia.org/P94606 and previous config saved to /var/cache/conftool/dbconfig/20260630-084632-fceratto.json
- 08:44 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 T430624', diff saved to https://phabricator.wikimedia.org/P94605 and previous config saved to /var/cache/conftool/dbconfig/20260630-084436-fceratto.json
- 08:40 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s2 T430624
- 08:39 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2099.codfw.wmnet with OS trixie
- 08:37 jmm@dns1004: END - running authdns-update
- 08:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1223', diff saved to https://phabricator.wikimedia.org/P94604 and previous config saved to /var/cache/conftool/dbconfig/20260630-083624-fceratto.json
- 08:35 jmm@dns1004: START - running authdns-update
- 08:34 filippo@dns1004: END - running authdns-update
- 08:32 filippo@dns1004: START - running authdns-update
- 08:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1223', diff saved to https://phabricator.wikimedia.org/P94603 and previous config saved to /var/cache/conftool/dbconfig/20260630-082616-fceratto.json
- 08:19 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2099.codfw.wmnet with reason: host reimage
- 08:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1223 (T426633)', diff saved to https://phabricator.wikimedia.org/P94602 and previous config saved to /var/cache/conftool/dbconfig/20260630-081609-fceratto.json
- 08:12 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2099.codfw.wmnet with reason: host reimage
- 08:08 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db1223 (T426633)', diff saved to https://phabricator.wikimedia.org/P94601 and previous config saved to /var/cache/conftool/dbconfig/20260630-080858-fceratto.json
- 08:08 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance
- 08:07 elukey: uploaded python3-wmflib_3.1.0 to apt.wikimedia.org bullseye-wikimedia,bookworm-wikimedia,trixie-wikimedia
- 07:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2098.codfw.wmnet with OS trixie
- 07:53 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2099.codfw.wmnet with OS trixie
- 07:45 moritzm: installing nodejs security updates
- 07:42 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1230 (T426633)', diff saved to https://phabricator.wikimedia.org/P94600 and previous config saved to /var/cache/conftool/dbconfig/20260630-074229-fceratto.json
- 07:40 moritzm: installing libtext-csv-xs-perl security updates
- 07:37 moritzm: installing libhttp-daemon-perl security updates
- 07:34 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2098.codfw.wmnet with reason: host reimage
- 07:32 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1230', diff saved to https://phabricator.wikimedia.org/P94599 and previous config saved to /var/cache/conftool/dbconfig/20260630-073221-fceratto.json
- 07:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2098.codfw.wmnet with reason: host reimage
- 07:23 moritzm: installing libgd-perl security updates
- 07:22 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1230', diff saved to https://phabricator.wikimedia.org/P94598 and previous config saved to /var/cache/conftool/dbconfig/20260630-072214-fceratto.json
- 07:21 elukey: upgrade all bullseye hosts to pywmflib 3.0 - T430552
- 07:12 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1230 (T426633)', diff saved to https://phabricator.wikimedia.org/P94597 and previous config saved to /var/cache/conftool/dbconfig/20260630-071206-fceratto.json
- 07:07 moritzm: installing libconfig-inifiles-perl security updates
- 07:07 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2098.codfw.wmnet with OS trixie
- 07:05 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db1230 (T426633)', diff saved to https://phabricator.wikimedia.org/P94596 and previous config saved to /var/cache/conftool/dbconfig/20260630-070512-fceratto.json
- 07:05 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance
- 06:56 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2064.codfw.wmnet with OS trixie
- 06:45 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Repooling after switchover
- 06:40 ryankemper@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cirrussearch2077.codfw.wmnet with OS trixie
- 06:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2064.codfw.wmnet with reason: host reimage
- 06:26 marostegui@dns1004: END - running authdns-update
- 06:25 ryankemper@cumin2002: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch2077.codfw.wmnet with reason: host reimage
- 06:22 marostegui@dns1004: START - running authdns-update
- 06:22 marostegui@dns1004: START - running authdns-update
- 06:16 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2064.codfw.wmnet with reason: host reimage
- 06:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2077.codfw.wmnet with reason: host reimage
- 06:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Repooling after switchover
- 06:00 marostegui@dns1004: END - running authdns-update
- 06:00 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Repooling after switchover
- 06:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1223: Repooling after switchover
- 05:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Repooling after switchover
- 05:59 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1223 T430610', diff saved to https://phabricator.wikimedia.org/P94588 and previous config saved to /var/cache/conftool/dbconfig/20260630-055906-marostegui.json
- 05:58 marostegui@dns1004: START - running authdns-update
- 05:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2077.codfw.wmnet with OS trixie
- 05:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2064.codfw.wmnet with OS trixie
- 05:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db1189 to s3 primary and set section read-write T430610', diff saved to https://phabricator.wikimedia.org/P94587 and previous config saved to /var/cache/conftool/dbconfig/20260630-055812-marostegui.json
- 05:57 marostegui@cumin1003: dbctl commit (dc=all): 'Set s3 eqiad as read-only for maintenance - T430610', diff saved to https://phabricator.wikimedia.org/P94586 and previous config saved to /var/cache/conftool/dbconfig/20260630-055752-marostegui.json
- 05:55 marostegui: Starting s3 eqiad failover from db1223 to db1189 - T430610
- 05:54 marostegui@cumin1003: dbctl commit (dc=all): 'Set db1189 with weight 0 T430610', diff saved to https://phabricator.wikimedia.org/P94585 and previous config saved to /var/cache/conftool/dbconfig/20260630-055423-marostegui.json
- 05:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 T430610
- 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2114.codfw.wmnet with OS trixie
- 05:30 marostegui@dns1004: END - running authdns-update
- 05:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Repooling after switchover
- 05:28 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1230: Repooling after switchover
- 05:28 marostegui@dns1004: START - running authdns-update
- 05:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Repooling after switchover
- 05:26 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1230 T430540', diff saved to https://phabricator.wikimedia.org/P94582 and previous config saved to /var/cache/conftool/dbconfig/20260630-052624-marostegui.json
- 05:25 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db1210 to s5 primary and set section read-write T430540', diff saved to https://phabricator.wikimedia.org/P94581 and previous config saved to /var/cache/conftool/dbconfig/20260630-052547-marostegui.json
- 05:25 marostegui@cumin1003: dbctl commit (dc=all): 'Set s5 eqiad as read-only for maintenance - T430540', diff saved to https://phabricator.wikimedia.org/P94580 and previous config saved to /var/cache/conftool/dbconfig/20260630-052523-marostegui.json
- 05:24 marostegui: Starting s5 eqiad failover from db1230 to db1210 - T430540
- 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 T430540
- 05:19 marostegui@cumin1003: dbctl commit (dc=all): 'Set db1210 with weight 0 T430540', diff saved to https://phabricator.wikimedia.org/P94579 and previous config saved to /var/cache/conftool/dbconfig/20260630-051920-marostegui.json
- 05:14 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2114.codfw.wmnet with reason: host reimage
- 05:07 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2114.codfw.wmnet with reason: host reimage
- 04:48 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2114.codfw.wmnet with OS trixie
- 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.6 (duration: 02m 36s)
- 03:58 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2239.codfw.wmnet with reason: Maintenance
- 03:58 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2227 (T410589)', diff saved to https://phabricator.wikimedia.org/P94578 and previous config saved to /var/cache/conftool/dbconfig/20260630-035825-ladsgroup.json
- 03:48 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2227', diff saved to https://phabricator.wikimedia.org/P94577 and previous config saved to /var/cache/conftool/dbconfig/20260630-034818-ladsgroup.json
- 03:41 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.9 refs T423918 (duration: 38m 39s)
- 03:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2227', diff saved to https://phabricator.wikimedia.org/P94576 and previous config saved to /var/cache/conftool/dbconfig/20260630-033809-ladsgroup.json
- 03:34 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2112.codfw.wmnet with OS trixie
- 03:28 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2227 (T410589)', diff saved to https://phabricator.wikimedia.org/P94575 and previous config saved to /var/cache/conftool/dbconfig/20260630-032802-ladsgroup.json
- 03:20 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2227 (T410589)', diff saved to https://phabricator.wikimedia.org/P94574 and previous config saved to /var/cache/conftool/dbconfig/20260630-032028-ladsgroup.json
- 03:20 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2227.codfw.wmnet with reason: Maintenance
- 03:20 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2209 (T410589)', diff saved to https://phabricator.wikimedia.org/P94573 and previous config saved to /var/cache/conftool/dbconfig/20260630-032004-ladsgroup.json
- 03:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2112.codfw.wmnet with reason: host reimage
- 03:09 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2209', diff saved to https://phabricator.wikimedia.org/P94572 and previous config saved to /var/cache/conftool/dbconfig/20260630-030956-ladsgroup.json
- 03:08 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2112.codfw.wmnet with reason: host reimage
- 03:06 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1163.eqiad.wmnet
- 03:06 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1163.eqiad.wmnet
- 03:06 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1163.eqiad.wmnet
- 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.9 refs T423918
- 02:59 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2209', diff saved to https://phabricator.wikimedia.org/P94571 and previous config saved to /var/cache/conftool/dbconfig/20260630-025948-ladsgroup.json
- 02:50 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2112.codfw.wmnet with OS trixie
- 02:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2209 (T410589)', diff saved to https://phabricator.wikimedia.org/P94570 and previous config saved to /var/cache/conftool/dbconfig/20260630-024941-ladsgroup.json
- 02:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2209 (T410589)', diff saved to https://phabricator.wikimedia.org/P94569 and previous config saved to /var/cache/conftool/dbconfig/20260630-024225-ladsgroup.json
- 02:42 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2209.codfw.wmnet with reason: Maintenance
- 02:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2194 (T410589)', diff saved to https://phabricator.wikimedia.org/P94568 and previous config saved to /var/cache/conftool/dbconfig/20260630-024212-ladsgroup.json
- 02:34 jasmine_: homer lsw1-d3-eqiad* commit 'T430226'
- 02:33 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1163.eqiad.wmnet with OS trixie
- 02:32 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2194', diff saved to https://phabricator.wikimedia.org/P94567 and previous config saved to /var/cache/conftool/dbconfig/20260630-023204-ladsgroup.json
- 02:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2194', diff saved to https://phabricator.wikimedia.org/P94566 and previous config saved to /var/cache/conftool/dbconfig/20260630-022157-ladsgroup.json
- 02:14 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1163.eqiad.wmnet with reason: host reimage
- 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2194 (T410589)', diff saved to https://phabricator.wikimedia.org/P94565 and previous config saved to /var/cache/conftool/dbconfig/20260630-021149-ladsgroup.json
- 02:09 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1163.eqiad.wmnet with reason: host reimage
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s)
- 02:04 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2194 (T410589)', diff saved to https://phabricator.wikimedia.org/P94564 and previous config saved to /var/cache/conftool/dbconfig/20260630-020416-ladsgroup.json
- 02:04 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2194.codfw.wmnet with reason: Maintenance
- 02:03 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2190 (T410589)', diff saved to https://phabricator.wikimedia.org/P94563 and previous config saved to /var/cache/conftool/dbconfig/20260630-020353-ladsgroup.json
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:54 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1163
- 01:54 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1163
- 01:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2190', diff saved to https://phabricator.wikimedia.org/P94562 and previous config saved to /var/cache/conftool/dbconfig/20260630-015345-ladsgroup.json
- 01:53 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1163
- 01:53 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1163.eqiad.wmnet 55.48.64.10.in-addr.arpa 5.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 01:52 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1163.eqiad.wmnet 55.48.64.10.in-addr.arpa 5.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 01:52 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 01:52 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1163 - jasmine@cumin2002"
- 01:52 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1163 - jasmine@cumin2002"
- 01:47 jasmine@cumin2002: START - Cookbook sre.dns.netbox
- 01:46 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1163
- 01:45 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1163.eqiad.wmnet with OS trixie
- 01:45 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1163.eqiad.wmnet
- 01:44 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1163.eqiad.wmnet
- 01:44 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1163.eqiad.wmnet
- 01:43 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2190', diff saved to https://phabricator.wikimedia.org/P94561 and previous config saved to /var/cache/conftool/dbconfig/20260630-014337-ladsgroup.json
- 01:37 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2063.codfw.wmnet with OS trixie
- 01:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2190 (T410589)', diff saved to https://phabricator.wikimedia.org/P94560 and previous config saved to /var/cache/conftool/dbconfig/20260630-013329-ladsgroup.json
- 01:26 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2190 (T410589)', diff saved to https://phabricator.wikimedia.org/P94559 and previous config saved to /var/cache/conftool/dbconfig/20260630-012605-ladsgroup.json
- 01:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2190.codfw.wmnet with reason: Maintenance
- 01:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2177 (T410589)', diff saved to https://phabricator.wikimedia.org/P94558 and previous config saved to /var/cache/conftool/dbconfig/20260630-012540-ladsgroup.json
- 01:18 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2063.codfw.wmnet with reason: host reimage
- 01:15 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2177', diff saved to https://phabricator.wikimedia.org/P94557 and previous config saved to /var/cache/conftool/dbconfig/20260630-011533-ladsgroup.json
- 01:12 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2063.codfw.wmnet with reason: host reimage
- 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2177', diff saved to https://phabricator.wikimedia.org/P94556 and previous config saved to /var/cache/conftool/dbconfig/20260630-010525-ladsgroup.json
- 00:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2177 (T410589)', diff saved to https://phabricator.wikimedia.org/P94555 and previous config saved to /var/cache/conftool/dbconfig/20260630-005517-ladsgroup.json
- 00:53 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2063.codfw.wmnet with OS trixie
- 00:47 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2177 (T410589)', diff saved to https://phabricator.wikimedia.org/P94554 and previous config saved to /var/cache/conftool/dbconfig/20260630-004721-ladsgroup.json
- 00:47 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2177.codfw.wmnet with reason: Maintenance
- 00:47 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2156 (T410589)', diff saved to https://phabricator.wikimedia.org/P94553 and previous config saved to /var/cache/conftool/dbconfig/20260630-004657-ladsgroup.json
- 00:36 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2156', diff saved to https://phabricator.wikimedia.org/P94552 and previous config saved to /var/cache/conftool/dbconfig/20260630-003650-ladsgroup.json
- 00:26 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2156', diff saved to https://phabricator.wikimedia.org/P94551 and previous config saved to /var/cache/conftool/dbconfig/20260630-002642-ladsgroup.json
- 00:16 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2156 (T410589)', diff saved to https://phabricator.wikimedia.org/P94550 and previous config saved to /var/cache/conftool/dbconfig/20260630-001634-ladsgroup.json
- 00:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2156 (T410589)', diff saved to https://phabricator.wikimedia.org/P94549 and previous config saved to /var/cache/conftool/dbconfig/20260630-000844-ladsgroup.json
- 00:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2156.codfw.wmnet with reason: Maintenance
2026-06-29
- 23:38 inflatador: bking@localhost raise the number of incoming shard recoveries from 4 to 7 on all search_codfw endpoints T429844
- 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cirrussearch2091.codfw.wmnet with OS trixie
- 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cirrussearch2089.codfw.wmnet']
- 21:21 maryum: Deployed security fix for T430548
- 21:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2073.codfw.wmnet with reason: host reimage
- 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2091.codfw.wmnet with reason: host reimage
- 21:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2073.codfw.wmnet with reason: host reimage
- 21:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cirrussearch2089.codfw.wmnet']
- 21:09 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2091.codfw.wmnet with reason: host reimage
- 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for SuggestedLinkEditCheck: by default import the config of the growth task (T422730), SuggestedLinkEditCheck: make non-experimental (T421968) (duration: 10m 25s)
- 20:59 kemayo@deploy1003: kemayo: Continuing with deployment
- 20:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2073.codfw.wmnet with OS trixie
- 20:55 kemayo@deploy1003: kemayo: Backport for SuggestedLinkEditCheck: by default import the config of the growth task (T422730), SuggestedLinkEditCheck: make non-experimental (T421968) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:53 kemayo@deploy1003: Started scap sync-world: Backport for SuggestedLinkEditCheck: by default import the config of the growth task (T422730), SuggestedLinkEditCheck: make non-experimental (T421968)
- 20:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2091.codfw.wmnet with OS trixie
- 20:39 kemayo@deploy1003: Finished scap sync-world: Backport for Add missing resolveUrlOrTitle helper function (T430450), Add missing visualeditor-suggestion-link message to extension.json (T430450) (duration: 07m 38s)
- 20:35 kemayo@deploy1003: kemayo: Continuing with deployment
- 20:33 kemayo@deploy1003: kemayo: Backport for Add missing resolveUrlOrTitle helper function (T430450), Add missing visualeditor-suggestion-link message to extension.json (T430450) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:31 kemayo@deploy1003: Started scap sync-world: Backport for Add missing resolveUrlOrTitle helper function (T430450), Add missing visualeditor-suggestion-link message to extension.json (T430450)
- 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid Read views for 5% of English Wikipedia desktop traffic (T430194), Temporarily disable experimental ExtTagPFragment type (T430344 T429624) (duration: 13m 16s)
- 20:18 arlolra@deploy1003: arlolra, cscott: Continuing with deployment
- 20:16 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye
- 20:11 arlolra@deploy1003: arlolra, cscott: Backport for Turn on Parsoid Read views for 5% of English Wikipedia desktop traffic (T430194), Temporarily disable experimental ExtTagPFragment type (T430344 T429624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:09 arlolra@deploy1003: Started scap sync-world: Backport for Turn on Parsoid Read views for 5% of English Wikipedia desktop traffic (T430194), Temporarily disable experimental ExtTagPFragment type (T430344 T429624)
- 19:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2090.codfw.wmnet with OS trixie
- 19:55 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:55 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:54 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:54 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:51 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 19:51 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 19:49 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 19:45 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2090.codfw.wmnet with reason: host reimage
- 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2075.codfw.wmnet with OS trixie
- 19:38 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply
- 19:38 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply
- 19:36 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply
- 19:35 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply
- 19:33 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply
- 19:33 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply
- 19:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2090.codfw.wmnet with reason: host reimage
- 19:27 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply
- 19:27 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply
- 19:22 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply
- 19:21 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply
- 19:19 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply
- 19:19 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply
- 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2075.codfw.wmnet with reason: host reimage
- 19:17 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply
- 19:16 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply
- 19:15 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply
- 19:15 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply
- 19:14 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply
- 19:14 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply
- 19:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2075.codfw.wmnet with reason: host reimage
- 19:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2090.codfw.wmnet with OS trixie
- 19:12 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply
- 19:11 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply
- 18:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2075.codfw.wmnet with OS trixie
- 18:53 zabe@deploy1003: Finished scap sync-world: Backport for Remove remaining occurences of apiportalwiki (T418494) (duration: 09m 11s)
- 18:48 zabe@deploy1003: zabe: Continuing with deployment
- 18:45 zabe@deploy1003: zabe: Backport for Remove remaining occurences of apiportalwiki (T418494) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:43 zabe@deploy1003: Started scap sync-world: Backport for Remove remaining occurences of apiportalwiki (T418494)
- 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2069.codfw.wmnet with OS trixie
- 18:39 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply
- 18:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply
- 18:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on contint2003.wikimedia.org with reason: not active yet
- 18:27 mutante: contint2003 - maintenance reboot
- 18:27 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on contint1003.wikimedia.org with reason: not active yet
- 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2069.codfw.wmnet with reason: host reimage
- 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2062.codfw.wmnet with OS trixie
- 18:17 mutante: contint1003 - maintenance reboot
- 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2069.codfw.wmnet with reason: host reimage
- 18:03 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply
- 18:02 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply
- 17:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2069.codfw.wmnet with OS trixie
- 17:55 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply
- 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2062.codfw.wmnet with reason: host reimage
- 17:55 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply
- 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2088.codfw.wmnet with OS trixie
- 17:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2062.codfw.wmnet with reason: host reimage
- 17:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2062.codfw.wmnet with OS trixie
- 17:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2088.codfw.wmnet with reason: host reimage
- 17:22 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2088.codfw.wmnet with reason: host reimage
- 17:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Migration of db1260.eqiad.wmnet completed
- 17:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cirrussearch2087.codfw.wmnet with OS trixie
- 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2088.codfw.wmnet with OS trixie
- 16:24 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Migration of db1260.eqiad.wmnet completed
- 16:06 jdrewniak@deploy1003: Finished scap sync-world: Backport for Assets build - 2026-06-29 15:36:06+00:00 (duration: 08m 37s)
- 16:04 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
- 15:58 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for Assets build - 2026-06-29 15:36:06+00:00 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:57 jdrewniak@deploy1003: Started scap sync-world: Backport for Assets build - 2026-06-29 15:36:06+00:00
- 15:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1260.eqiad.wmnet with OS trixie
- 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2087.codfw.wmnet with reason: host reimage
- 15:51 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 15:50 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
- 15:50 moritzm: installing libconfig-inifiles-perl security updates
- 15:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2087.codfw.wmnet with reason: host reimage
- 15:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cirrussearch2074.codfw.wmnet with OS trixie
- 15:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2074.codfw.wmnet with reason: host reimage
- 15:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1260.eqiad.wmnet with reason: host reimage
- 15:38 moritzm: installing zsh updates from Trixie point release
- 15:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2074.codfw.wmnet with reason: host reimage
- 15:34 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1260.eqiad.wmnet with reason: host reimage
- 15:32 moritzm: installing glib2.0 security updates
- 15:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2087.codfw.wmnet with OS trixie
- 15:24 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply
- 15:24 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply
- 15:19 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1260.eqiad.wmnet with OS trixie
- 15:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2074.codfw.wmnet with OS trixie
- 15:14 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply
- 15:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply
- 15:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
- 14:48 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [config] Fix code in core-Permissions.php (duration: 07m 03s)
- 14:43 dreamyjazz@deploy1003: vadymts1, dreamyjazz: Continuing with deployment
- 14:42 dreamyjazz@deploy1003: vadymts1, dreamyjazz: Backport for [config] Fix code in core-Permissions.php synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [config] Fix code in core-Permissions.php
- 14:25 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
- 14:18 kharlan@deploy1003: Finished scap sync-world: Backport for hCaptcha: Align the loginattempt CAPTCHA with badlogin (T428892) (duration: 07m 25s)
- 14:14 kharlan@deploy1003: kharlan: Continuing with deployment
- 14:12 kharlan@deploy1003: kharlan: Backport for hCaptcha: Align the loginattempt CAPTCHA with badlogin (T428892) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:11 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
- 14:10 kharlan@deploy1003: Started scap sync-world: Backport for hCaptcha: Align the loginattempt CAPTCHA with badlogin (T428892)
- 14:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Upgrading db1260.eqiad.wmnet
- 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Upgrading db1260.eqiad.wmnet
- 14:07 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 14:05 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 13:57 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
- 13:55 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 13:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
- 13:52 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
- 13:51 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
- 13:49 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
- 13:49 Tran: Deployed patch for T427287
- 13:39 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 13:33 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 13:31 stran@deploy1003: Scap cancelled without rolling back.
- 13:31 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
- 13:27 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
- 13:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 13:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Migration of db1252.eqiad.wmnet completed
- 13:15 stran@deploy1003: stran, vadymts1: Continuing with deployment
- 13:13 stran@deploy1003: stran, vadymts1: Backport for User groups changes for English Wikiversity (T430416), hrwiki: Add to wgCiteResponsiveReferences (T430182) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:11 stran@deploy1003: Started scap sync-world: Backport for User groups changes for English Wikiversity (T430416), hrwiki: Add to wgCiteResponsiveReferences (T430182)
- 13:05 dpogorzelski@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: sync
- 13:05 dpogorzelski@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: sync
- 13:04 dpogorzelski@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: sync
- 13:04 dpogorzelski@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: sync
- 12:46 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
- 12:46 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
- 12:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
- 12:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
- 12:41 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
- 12:41 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
- 12:37 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Migration of db1252.eqiad.wmnet completed
- 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
- 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
- 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
- 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
- 12:33 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
- 12:33 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
- 12:32 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
- 12:32 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
- 12:29 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
- 12:29 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
- 12:28 godog: add cloudvirt10[78-80] to nova -- with compute disabled - T429563
- 12:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1252.eqiad.wmnet with OS trixie
- 12:19 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
- 12:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
- 12:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
- 12:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
- 12:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
- 12:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
- 12:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
- 12:11 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
- 12:11 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
- 12:10 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
- 12:10 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
- 12:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
- 12:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
- 12:07 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 12:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1252.eqiad.wmnet with reason: host reimage
- 12:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
- 12:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
- 11:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 11:57 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1252.eqiad.wmnet with reason: host reimage
- 11:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 11:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply}
- 11:55 cmooney@dns3003: END - running authdns-update
- 11:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
- 11:53 cmooney@dns3003: START - running authdns-update
- 11:47 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:47 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for private1-d8test-eqiad vlan - cmooney@cumin1003"
- 11:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for private1-d8test-eqiad vlan - cmooney@cumin1003"
- 11:43 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 11:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 11:40 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1252.eqiad.wmnet with OS trixie
- 11:40 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 11:39 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
- 11:39 cgoubert@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
- 11:38 cgoubert@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
- 11:38 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1006.eqiad.wmnet with OS trixie
- 11:38 cgoubert@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
- 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Upgrading db1252.eqiad.wmnet
- 11:37 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Upgrading db1252.eqiad.wmnet
- 11:36 cgoubert@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
- 11:36 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 11:36 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 11:36 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 11:34 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
- 11:33 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 11:32 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
- 11:31 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
- 11:30 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
- 11:30 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 11:29 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 11:28 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 11:27 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 11:27 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 11:26 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 11:19 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1006.eqiad.wmnet with reason: host reimage
- 11:13 cmooney@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1006.eqiad.wmnet with reason: host reimage
- 10:55 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 10:54 cmooney@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest1006.eqiad.wmnet with OS trixie
- 10:48 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 10:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2247: Migration of db2247.codfw.wmnet completed
- 10:47 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 10:41 cmooney@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest1006.eqiad.wmnet with OS trixie
- 10:35 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Migration of db1249.eqiad.wmnet completed
- 10:28 aikochou@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
- 10:23 aikochou@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 10:17 aikochou@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 10:02 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2247: Migration of db2247.codfw.wmnet completed
- 09:54 moritzm: installing libssh2 security updates
- 09:49 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Migration of db1249.eqiad.wmnet completed
- 09:46 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2247.codfw.wmnet with OS trixie
- 09:39 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw
- 09:36 marostegui: drop database apiportalwiki T418494
- 09:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1249.eqiad.wmnet with OS trixie
- 09:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2247.codfw.wmnet with reason: host reimage
- 09:25 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2247.codfw.wmnet with reason: host reimage
- 09:22 dcausse: T418494 delete apiportalwiki cirrussearch indices
- 09:19 moritzm: installing librabbitmq security updates
- 09:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1249.eqiad.wmnet with reason: host reimage
- 09:13 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1249.eqiad.wmnet with reason: host reimage
- 09:12 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 09:12 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 09:10 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2247.codfw.wmnet with OS trixie
- 09:09 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 09:00 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs-test2001.codfw.wmnet
- 08:59 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-test-wdqs2001.codfw.wmnet
- 08:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1080.eqiad.wmnet
- 08:59 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-test-wdqs2001.codfw.wmnet
- 08:58 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2003.codfw.wmnet
- 08:55 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Upgrading db2247.codfw.wmnet
- 08:55 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Upgrading db2247.codfw.wmnet
- 08:55 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 08:54 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1249.eqiad.wmnet with OS trixie
- 08:52 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1249: Upgrading db1249.eqiad.wmnet
- 08:51 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1249: Upgrading db1249.eqiad.wmnet
- 08:51 Tran: Deployed patch for T427287
- 08:51 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 08:51 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1079.eqiad.wmnet
- 08:43 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1078.eqiad.wmnet
- 08:40 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
- 08:34 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1080.eqiad.wmnet
- 08:34 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1079.eqiad.wmnet
- 08:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1078.eqiad.wmnet
- 08:32 moritzm: installing libxslt bugfix updates
- 08:25 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 08:09 moritzm: installing lcms2 security updates
- 08:07 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
- 08:06 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
- 08:05 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
- 08:04 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
- 08:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
- 08:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
- 07:45 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1026.eqiad.wmnet
- 07:23 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 07:23 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- 07:23 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 07:23 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
- 07:02 tappof: bump space for prometheus k8s-dse in eqiad
- 06:04 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2021: pc1 repool
- 06:04 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 06:04 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 06:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2021: pc1 repool
- 05:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on pc2021.codfw.wmnet,pc1021.eqiad.wmnet with reason: Debugging
- 05:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2021: pc1 hw issues
- 05:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 05:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 05:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2021: pc1 hw issues
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 05s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-28
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-27
- 06:44 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 06:44 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 06:44 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 06:44 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 05:42 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 05:41 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 05:41 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 05:41 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-26
- 21:18 denisse: Upgrading Grafana in grafana2001 - T430112
- 21:00 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host restbase2039.codfw.wmnet with OS bullseye
- 20:24 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye
- 19:18 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:07 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 18:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 18:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new esams arelion cct - cmooney@cumin1003"
- 18:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new esams arelion cct - cmooney@cumin1003"
- 18:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 17:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1023.eqiad.wmnet
- 17:41 topranks: enable "graceful shutdown" community on cr2-eqiad to allow for reset of line card 1/1 T427843
- 17:35 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1023.eqiad.wmnet
- 17:27 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 17:26 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 17:26 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 17:25 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 17:24 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 17:24 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add esams router ips - cmooney@cumin1003"
- 17:24 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add esams router ips - cmooney@cumin1003"
- 17:20 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 17:20 cmooney@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
- 17:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 16:44 dwisehaupt@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:41 dwisehaupt@cumin1003: START - Cookbook sre.dns.netbox
- 16:38 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 16:33 cmooney@dns2005: END - running authdns-update
- 16:31 cmooney@dns2005: START - running authdns-update
- 16:24 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:24 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for esams transport IPs - cmooney@cumin1003"
- 16:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for esams transport IPs - cmooney@cumin1003"
- 16:19 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1023.eqiad.wmnet with reason: host reimage
- 16:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1023.eqiad.wmnet with reason: host reimage
- 15:56 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 15:35 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 15:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1023.eqiad.wmnet with reason: host reimage
- 15:19 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1023.eqiad.wmnet with reason: host reimage
- 15:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 15:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 15:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 15:02 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 14:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
- 14:50 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
- 14:46 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1023.eqiad.wmnet with OS bookworm
- 14:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host dse-k8s-worker1023.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 14:37 btullis@cumin1003: START - Cookbook sre.hosts.provision for host dse-k8s-worker1023.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 14:33 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
- 14:32 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
- 14:32 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
- 14:32 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
- 14:29 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
- 14:29 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
- 14:25 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
- 14:25 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
- 14:04 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023
- 14:04 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023
- 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Moved dse-k8s-worker1023 vlan - btullis@cumin1003"
- 14:02 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Moved dse-k8s-worker1023 vlan - btullis@cumin1003"
- 13:59 krinkle@deploy1003: Finished scap sync-world: Backport for Undeploy the ShortUrl extension (T107188), extension-list: Remove WikimediaApiPortalOAuth ext and WikimediaApiPortal skin (T429373 T429374 T418494) (duration: 44m 35s)
- 13:56 btullis@cumin1003: START - Cookbook sre.dns.netbox
- 13:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts dse-k8s-worker1023.eqiad.wmnet
- 13:48 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 13:48 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: dse-k8s-worker1023.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
- 13:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: dse-k8s-worker1023.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
- 13:46 krinkle@deploy1003: krinkle: Continuing with deployment
- 13:43 btullis@cumin1003: START - Cookbook sre.dns.netbox
- 13:36 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts dse-k8s-worker1023.eqiad.wmnet
- 13:32 krinkle@deploy1003: krinkle: Backport for Undeploy the ShortUrl extension (T107188), extension-list: Remove WikimediaApiPortalOAuth ext and WikimediaApiPortal skin (T429373 T429374 T418494) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 13:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 13:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
- 13:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
- 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
- 13:14 krinkle@deploy1003: Started scap sync-world: Backport for Undeploy the ShortUrl extension (T107188), extension-list: Remove WikimediaApiPortalOAuth ext and WikimediaApiPortal skin (T429373 T429374 T418494)
- 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
- 13:09 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 13:05 hashar: Restarting CI Jenkins
- 12:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
- 12:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
- 12:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
- 12:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1026.eqiad.wmnet,service=s1
- 12:32 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1026.eqiad.wmnet
- 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
- 12:30 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1026.eqiad.wmnet,service=s1
- 12:30 marostegui@cumin1003: conftool action : set/weight=1; selector: name=clouddb1026.eqiad.wmnet
- 12:24 oblivian@cumin1003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool zotero in eqiad: maintenance
- 12:19 oblivian@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) zotero.discovery.wmnet on all recursors
- 12:19 oblivian@cumin1003: START - Cookbook sre.dns.wipe-cache zotero.discovery.wmnet on all recursors
- 12:19 oblivian@cumin1003: START - Cookbook sre.discovery.service-route pool zotero in eqiad: maintenance
- 12:19 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
- 12:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
- 12:17 oblivian@cumin1003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check citoid: maintenance
- 12:17 oblivian@cumin1003: START - Cookbook sre.discovery.service-route check citoid: maintenance
- 12:15 oblivian@cumin1003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) depool zotero in eqiad: Testing theory about upstreams
- 12:10 oblivian@cumin1003: START - Cookbook sre.discovery.service-route depool zotero in eqiad: Testing theory about upstreams
- 12:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
- 12:05 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
- 12:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
- 12:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 12:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
- 11:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 11:43 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: sync
- 11:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 11:42 _joe_: rolling restart of citoid in eqiad T430279
- 11:42 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: sync
- 11:32 _joe_: rolling restart of zotero in eqiad T430279
- 11:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: sync
- 11:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: sync
- 11:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1080.eqiad.wmnet with OS trixie
- 11:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 11:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 11:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1079.eqiad.wmnet with OS trixie
- 11:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
- 11:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1078.eqiad.wmnet with OS trixie
- 11:13 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1080.eqiad.wmnet with reason: host reimage
- 11:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1079.eqiad.wmnet with reason: host reimage
- 11:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1078.eqiad.wmnet with reason: host reimage
- 10:57 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1080.eqiad.wmnet with reason: host reimage
- 10:57 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1079.eqiad.wmnet with reason: host reimage
- 10:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1078.eqiad.wmnet with reason: host reimage
- 10:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 10:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 10:45 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1080.eqiad.wmnet with OS trixie
- 10:45 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1079.eqiad.wmnet with OS trixie
- 10:45 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1078.eqiad.wmnet with OS trixie
- 10:44 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 10:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new IPs for cloudvirt1078-80 - filippo@cumin1003"
- 10:44 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new IPs for cloudvirt1078-80 - filippo@cumin1003"
- 10:40 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 10:12 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database isvwiki (T429938)
- 09:42 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database magwiki (T428282)
- 09:42 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database magwiki (T428282)
- 09:15 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database isvwiki (T429938)
- 09:09 fnegri@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.add-wiki (exit_code=99) for database isvwiki (T429938)
- 09:09 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database isvwiki (T429938)
- 09:05 fnegri@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.add-wiki (exit_code=99) for database isvwiki (T429938)
- 09:05 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database isvwiki (T429938)
- 07:39 jelto@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-unattended (exit_code=99) Sequential unattended reboot of 6 host(s) [team=collaboration-services, os=bookworm]
- 07:38 jelto@cumin1003: START - Cookbook sre.hosts.reboot-unattended Sequential unattended reboot of 6 host(s) [team=collaboration-services, os=bookworm]
- 05:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2234.codfw.wmnet with OS trixie
- 05:28 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2234.codfw.wmnet with reason: host reimage
- 05:23 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2234.codfw.wmnet with reason: host reimage
- 05:08 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2234.codfw.wmnet with OS trixie
- 04:08 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
- 04:08 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 09s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:17 ryankemper: T429919 Followed the steps in https://wikitech.wikimedia.org/wiki/Wikidata_Query_Service/Streaming_Updater#Clean_up_object_storage to clear out stale checkpoints
2026-06-25
- 23:56 ryankemper@cumin2002: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
- 23:41 ryankemper@cumin2002: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
- 23:31 ryankemper@cumin2002: START - Cookbook sre.wdqs.restart
- 23:31 ryankemper@cumin2002: START - Cookbook sre.wdqs.restart
- 21:56 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2003
- 21:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for Restore carousel beta opt-in, suppressed when enabled sitewide (T429414), Roll back mobile image carousel from sitewide to beta opt-in (T429414) (duration: 32m 28s)
- 21:35 jdlrobson@deploy1003: egardner, jdlrobson: Continuing with deployment
- 21:33 jdlrobson@deploy1003: egardner, jdlrobson: Backport for Restore carousel beta opt-in, suppressed when enabled sitewide (T429414), Roll back mobile image carousel from sitewide to beta opt-in (T429414) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2003
- 21:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 21:15 jdlrobson@deploy1003: Started scap sync-world: Backport for Restore carousel beta opt-in, suppressed when enabled sitewide (T429414), Roll back mobile image carousel from sitewide to beta opt-in (T429414)
- 21:11 dani@deploy1003: Finished scap sync-world: Backport for Undeploy English Wikipedia Mobile App Survey (T428876) (duration: 06m 54s)
- 21:06 dani@deploy1003: dani: Continuing with deployment
- 21:06 dani@deploy1003: dani: Backport for Undeploy English Wikipedia Mobile App Survey (T428876) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:04 dani@deploy1003: Started scap sync-world: Backport for Undeploy English Wikipedia Mobile App Survey (T428876)
- 20:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 20:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 20:55 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 20:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 20:51 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 20:17 arlolra@deploy1003: Finished scap sync-world: Backport for Deploy PRV to 5 wikis (T429830), Enable ULS v2 by default across all wikis, Remove wgCiteRemoveSyntheticRefsUnsafe feature flag from production and beta cluster config (T428232) (duration: 08m 04s)
- 20:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2045: rack depool
- 20:12 arlolra@deploy1003: arlolra, mareikeheuer, abi: Continuing with deployment
- 20:10 arlolra@deploy1003: arlolra, mareikeheuer, abi: Backport for Deploy PRV to 5 wikis (T429830), Enable ULS v2 by default across all wikis, Remove wgCiteRemoveSyntheticRefsUnsafe feature flag from production and beta cluster config (T428232) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:08 arlolra@deploy1003: Started scap sync-world: Backport for Deploy PRV to 5 wikis (T429830), Enable ULS v2 by default across all wikis, Remove wgCiteRemoveSyntheticRefsUnsafe feature flag from production and beta cluster config (T428232)
- 19:59 cscott@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (take 2) (T429624 T429822 T391624) (duration: 06m 58s)
- 19:54 cscott@deploy1003: cscott: Continuing with deployment
- 19:53 cscott@deploy1003: cscott: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (take 2) (T429624 T429822 T391624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 19:52 cscott@deploy1003: Started scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (take 2) (T429624 T429822 T391624)
- 19:49 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1162.eqiad.wmnet with OS trixie
- 19:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1162.eqiad.wmnet with reason: host reimage
- 19:27 cscott@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624) (duration: 22m 07s)
- 19:27 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2045: rack depool
- 19:26 cscott@deploy1003: Rolling back deployment
- 19:24 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1162.eqiad.wmnet with reason: host reimage
- 19:18 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 19:12 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
- 19:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004
- 19:11 cscott@deploy1003: cscott, arlolra: Continuing with deployment
- 19:10 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004
- 19:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:08 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2003
- 19:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2003
- 19:08 bking@cumin2002: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:08 bking@cumin2002: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:08 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1162
- 19:08 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1162
- 19:07 bking@cumin2002: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:07 bking@cumin2002: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:07 cscott@deploy1003: cscott, arlolra: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 19:07 bking@cumin2002: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:07 bking@cumin2002: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: T426862 - bking@cumin2002
- 19:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 19:06 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1162
- 19:06 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1162.eqiad.wmnet 108.48.64.10.in-addr.arpa 8.0.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 19:06 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1162.eqiad.wmnet 108.48.64.10.in-addr.arpa 8.0.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 19:06 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:05 cscott@deploy1003: Started scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624)
- 19:03 jasmine@cumin2002: START - Cookbook sre.dns.netbox
- 19:03 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:03 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
- 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
- 18:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 18:57 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1162
- 18:57 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 18:57 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
- 18:57 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
- 18:57 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1162.eqiad.wmnet with OS trixie
- 18:50 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 18:45 reedy@deploy1003: Finished scap sync-world: Backport for InitialiseSettings: Require 2FA for all on arbcom_*wiki and conductwiki (T428103) (duration: 07m 12s)
- 18:43 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004
- 18:41 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1161.eqiad.wmnet with OS trixie
- 18:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004
- 18:41 reedy@deploy1003: reedy: Continuing with deployment
- 18:40 reedy@deploy1003: reedy: Backport for InitialiseSettings: Require 2FA for all on arbcom_*wiki and conductwiki (T428103) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:38 reedy@deploy1003: Started scap sync-world: Backport for InitialiseSettings: Require 2FA for all on arbcom_*wiki and conductwiki (T428103)
- 18:28 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.8 refs T423917
- 18:26 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004
- 18:25 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004
- 18:23 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 18:20 vriley@cumin1003: START - Cookbook sre.dns.netbox
- 18:20 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1161.eqiad.wmnet with reason: host reimage
- 18:15 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1161.eqiad.wmnet with reason: host reimage
- 18:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 18:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2246: Migration of db2246.codfw.wmnet completed
- 18:03 reedy@deploy1003: Synchronized private/: (no justification provided) (duration: 06m 14s)
- 18:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 18:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Migration of db1248.eqiad.wmnet completed
- 17:58 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1161
- 17:58 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1161
- 17:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 17:47 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1161
- 17:47 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1161.eqiad.wmnet 118.48.64.10.in-addr.arpa 8.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 17:47 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1161.eqiad.wmnet 118.48.64.10.in-addr.arpa 8.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 17:47 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 17:47 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1161 - jasmine@cumin2002"
- 17:47 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1161 - jasmine@cumin2002"
- 17:42 cscott@deploy1003: Finished scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624) (duration: 51m 59s)
- 17:41 cscott@deploy1003: Rolling back deployment
- 17:38 jasmine@cumin2002: START - Cookbook sre.dns.netbox
- 17:38 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1161
- 17:37 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1161.eqiad.wmnet with OS trixie
- 17:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
- 17:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
- 17:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
- 17:33 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
- 17:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
- 17:32 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
- 17:30 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1160.eqiad.wmnet with OS trixie
- 17:22 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2246: Migration of db2246.codfw.wmnet completed
- 17:15 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Migration of db1248.eqiad.wmnet completed
- 17:09 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1160.eqiad.wmnet with reason: host reimage
- 17:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2246.codfw.wmnet with OS trixie
- 17:04 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1160.eqiad.wmnet with reason: host reimage
- 17:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1248.eqiad.wmnet with OS trixie
- 16:59 cscott@deploy1003: cscott: Continuing with deployment
- 16:52 cscott@deploy1003: cscott: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 16:50 cscott@deploy1003: Started scap sync-world: Backport for Turn on Parsoid's 'ReturnExperimentalPFragmentTypes' (T429624 T429822 T391624)
- 16:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2246.codfw.wmnet with reason: host reimage
- 16:48 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1160
- 16:48 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1160
- 16:47 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1160
- 16:47 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1160.eqiad.wmnet 116.48.64.10.in-addr.arpa 6.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 16:47 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1160.eqiad.wmnet 116.48.64.10.in-addr.arpa 6.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 16:47 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:47 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1160 - jasmine@cumin2002"
- 16:47 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1160 - jasmine@cumin2002"
- 16:46 cscott@deploy1003: Finished scap sync-world: Backport for Bump wikimedia/parsoid to 0.24.0-a12 (T353697 T384490 T387374 T387520 T387521 T391624 T393295 T420336 T429624 T429688 T429822), Bump wikimedia/parsoid to 0.24.0-a12 (T429822) (duration: 09m 54s)
- 16:45 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1248.eqiad.wmnet with reason: host reimage
- 16:42 jasmine@cumin2002: START - Cookbook sre.dns.netbox
- 16:42 cscott@deploy1003: cscott: Continuing with deployment
- 16:41 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1160
- 16:41 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2246.codfw.wmnet with reason: host reimage
- 16:41 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1160.eqiad.wmnet with OS trixie
- 16:40 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1248.eqiad.wmnet with reason: host reimage
- 16:39 cscott@deploy1003: cscott: Backport for Bump wikimedia/parsoid to 0.24.0-a12 (T353697 T384490 T387374 T387520 T387521 T391624 T393295 T420336 T429624 T429688 T429822), Bump wikimedia/parsoid to 0.24.0-a12 (T429822) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 16:37 cscott@deploy1003: Started scap sync-world: Backport for Bump wikimedia/parsoid to 0.24.0-a12 (T353697 T384490 T387374 T387520 T387521 T391624 T393295 T420336 T429624 T429688 T429822), Bump wikimedia/parsoid to 0.24.0-a12 (T429822)
- 16:25 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2246.codfw.wmnet with OS trixie
- 16:23 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1248.eqiad.wmnet with OS trixie
- 16:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1248: Upgrading db1248.eqiad.wmnet
- 16:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Upgrading db2246.codfw.wmnet
- 16:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Upgrading db2246.codfw.wmnet
- 16:20 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 16:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 16:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1248: Upgrading db1248.eqiad.wmnet
- 16:20 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 16:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 16:17 pt1979@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cr2-eqdfw,cr2-eqdfw IPv6
- 16:17 pt1979@cumin2003: START - Cookbook sre.hosts.remove-downtime for cr2-eqdfw,cr2-eqdfw IPv6
- 16:02 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2234.codfw.wmnet with OS trixie
- 15:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2245: Migration of db2245.codfw.wmnet completed
- 15:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Migration of db1247.eqiad.wmnet completed
- 15:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 15:47 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 15:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 15:27 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 15:26 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 15:26 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 15:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 15:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 15:25 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 15:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2245: Migration of db2245.codfw.wmnet completed
- 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host build2004.codfw.wmnet with OS trixie
- 15:06 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Migration of db1247.eqiad.wmnet completed
- 15:02 pt1979@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-eqdfw,cr2-eqdfw IPv6 with reason: junos upgrade
- 15:00 pt1979@cumin2003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cr2-eqdfw cr2-eqdfw IPv6 with reason: junos upgrade
- 14:59 papaul: ongoing maintenance on cr2-eqdfw
- 14:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2245.codfw.wmnet with OS trixie
- 14:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on build2004.codfw.wmnet with reason: host reimage
- 14:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1247.eqiad.wmnet with OS trixie
- 14:48 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on build2004.codfw.wmnet with reason: host reimage
- 14:42 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2234.codfw.wmnet with OS trixie
- 14:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2245.codfw.wmnet with reason: host reimage
- 14:36 marostegui: Drop database apiportalwiki on sanitarium and wikireplicas T430102
- 14:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1247.eqiad.wmnet with reason: host reimage
- 14:30 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2245.codfw.wmnet with reason: host reimage
- 14:28 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1247.eqiad.wmnet with reason: host reimage
- 14:28 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host build2004.codfw.wmnet with OS trixie
- 14:21 hashar: Restarting CI Jenkins on contint1002
- 14:14 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2245.codfw.wmnet with OS trixie
- 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Upgrading db2245.codfw.wmnet
- 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Upgrading db2245.codfw.wmnet
- 14:11 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 14:11 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 14:11 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1247.eqiad.wmnet with OS trixie
- 14:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1247: Upgrading db1247.eqiad.wmnet
- 14:08 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1247: Upgrading db1247.eqiad.wmnet
- 14:08 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 14:08 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 14:05 Dreamy_Jazz: Ran `delete from cuci_user where ciu_ciwm_id = 4;` for T430156
- 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1077.eqiad.wmnet with OS trixie
- 14:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Migration of db2237.codfw.wmnet completed
- 13:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 13:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Migration of db1243.eqiad.wmnet completed
- 13:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2224: rack depool
- 13:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: rack depool
- 13:51 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host build2004.codfw.wmnet
- 13:51 jmm@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host build2004.codfw.wmnet with OS trixie
- 13:37 moritzm: installing imagemagick security updates
- 13:28 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 13:28 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 13:28 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 13:28 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 13:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2028.codfw.wmnet to cluster codfw and group A
- 13:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1077.eqiad.wmnet with reason: host reimage
- 13:16 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 13:16 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 13:16 zabe: zabe@deploy1003:~$ mwscript namespaceDupes.php isvwiki --fix # T429935
- 13:15 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Migration of db2237.codfw.wmnet completed
- 13:14 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 13:14 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 13:14 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 13:13 zabe@deploy1003: Finished scap sync-world: Backport for isvwiki: set timezone, sitename and logos (T429935), Use Hadoop for Mostcategories on commonswiki (T413362) (duration: 07m 30s)
- 13:13 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 13:12 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1077.eqiad.wmnet with reason: host reimage
- 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Migration of db1243.eqiad.wmnet completed
- 13:09 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2224: rack depool
- 13:09 zabe@deploy1003: zabe, anzx: Continuing with deployment
- 13:09 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2225: rack depool
- 13:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet,dse-k8s-wdqs-test2001.codfw.wmnet
- 13:08 zabe@deploy1003: zabe, anzx: Backport for isvwiki: set timezone, sitename and logos (T429935), Use Hadoop for Mostcategories on commonswiki (T413362) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:08 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet,dse-k8s-wdqs-test2001.codfw.wmnet
- 13:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2258-2259].codfw.wmnet
- 13:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2258-2259].codfw.wmnet
- 13:06 moritzm: re-added ganeti2028 to codfw/A Ganeti cluster T429817
- 13:06 zabe@deploy1003: Started scap sync-world: Backport for isvwiki: set timezone, sitename and logos (T429935), Use Hadoop for Mostcategories on commonswiki (T413362)
- 13:05 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2028.codfw.wmnet to cluster codfw and group A
- 13:02 moritzm: installing glib2.0 security updates
- 13:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1077.eqiad.wmnet with OS trixie
- 12:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2237.codfw.wmnet with OS trixie
- 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1243.eqiad.wmnet with OS trixie
- 12:44 XioNoX: lsw1-a7-codfw> request system reboot - T429817
- 12:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2237.codfw.wmnet with reason: host reimage
- 12:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1243.eqiad.wmnet with reason: host reimage
- 12:35 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2237.codfw.wmnet with reason: host reimage
- 12:33 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1243.eqiad.wmnet with reason: host reimage
- 12:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 31 hosts with reason: Rack A7 depool
- 12:04 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack A7
- 11:57 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-a7-codfw,lsw1-a7-codfw IPv6,lsw1-a7-codfw.mgmt with reason: Switch maintenance
- 11:53 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on build2004.codfw.wmnet with reason: host reimage
- 11:50 moritzm: installing harfbuzz security updates
- 11:48 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on build2004.codfw.wmnet with reason: host reimage
- 11:28 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host build2004.codfw.wmnet with OS trixie
- 11:26 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM build2004.codfw.wmnet - jmm@cumin2003"
- 11:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM build2004.codfw.wmnet - jmm@cumin2003"
- 11:25 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) build2004.codfw.wmnet on all recursors
- 11:25 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache build2004.codfw.wmnet on all recursors
- 11:25 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:25 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM build2004.codfw.wmnet - jmm@cumin2003"
- 11:25 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM build2004.codfw.wmnet - jmm@cumin2003"
- 11:14 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 11:14 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 11:09 jmm@cumin2003: START - Cookbook sre.dns.netbox
- 11:09 jmm@cumin2003: START - Cookbook sre.ganeti.makevm for new host build2004.codfw.wmnet
- 11:05 James_F: jforrester@deploy1003: mwscript sql.php --wiki=wikifunctionswiki --cluster extension1 extensions/WikiLambda/sql/mysql/table-wikifunctions_usage_wikis.sql # T428667
- 11:05 James_F: jforrester@deploy1003: mwscript sql.php --wiki=wikifunctionswiki --cluster extension1 extensions/WikiLambda/sql/mysql/table-wikifunctions_usage.sql # T428667
- 10:38 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host build2003.codfw.wmnet
- 10:38 jmm@cumin2002: START - Cookbook sre.ganeti.makevm for new host build2003.codfw.wmnet
- 10:37 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host build2003.codfw.wmnet
- 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.makevm for new host build2003.codfw.wmnet
- 10:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Migration of db2236.codfw.wmnet completed
- 10:28 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:28 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Migration of db1221.eqiad.wmnet completed
- 10:16 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1026.eqiad.wmnet,service=s1
- 10:15 blake@deploy1003: Stopping before sync operations
- 10:15 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values
- 10:14 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-staging-master@codfw
- 10:14 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 10:13 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
- 10:09 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-staging-master@codfw
- 10:07 klausman@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: ml-staging-master@codfw
- 10:03 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-staging-master@codfw
- 09:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging2008.codfw.wmnet with OS trixie
- 09:58 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:58 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging2007.codfw.wmnet with OS trixie
- 09:58 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:53 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 6648
- 09:52 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 6648
- 09:45 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Migration of db2236.codfw.wmnet completed
- 09:44 kharlan@deploy1003: Finished scap sync-world: Backport for hCaptcha: Skip blocked-IP score collection for crawlers in VE and MF (T429755) (duration: 08m 46s)
- 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Migration of db1221.eqiad.wmnet completed
- 09:40 kharlan@deploy1003: kharlan: Continuing with deployment
- 09:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging2006.codfw.wmnet with OS trixie
- 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging2008.codfw.wmnet with reason: host reimage
- 09:37 kharlan@deploy1003: kharlan: Backport for hCaptcha: Skip blocked-IP score collection for crawlers in VE and MF (T429755) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:35 kharlan@deploy1003: Started scap sync-world: Backport for hCaptcha: Skip blocked-IP score collection for crawlers in VE and MF (T429755)
- 09:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging2007.codfw.wmnet with reason: host reimage
- 09:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2236.codfw.wmnet with OS trixie
- 09:31 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging2008.codfw.wmnet with reason: host reimage
- 09:31 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging2007.codfw.wmnet with reason: host reimage
- 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1221.eqiad.wmnet with OS trixie
- 09:26 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2236.codfw.wmnet with reason: host reimage
- 09:13 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging2008.codfw.wmnet with OS trixie
- 09:12 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging2007.codfw.wmnet with OS trixie
- 09:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1221.eqiad.wmnet with reason: host reimage
- 09:11 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2236.codfw.wmnet with reason: host reimage
- 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging2006.codfw.wmnet with reason: host reimage
- 09:06 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1026.eqiad.wmnet
- 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1221.eqiad.wmnet with reason: host reimage
- 09:05 jforrester@deploy1003: Finished scap sync-world: Backport for On AW article deletion, clear all AWArticleStore from sections and metadata (T429873), AWStorage: Use global stash keys (T430060) (duration: 07m 29s)
- 09:05 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging2006.codfw.wmnet with reason: host reimage
- 09:00 jforrester@deploy1003: jforrester: Continuing with deployment
- 09:00 jforrester@deploy1003: jforrester: Backport for On AW article deletion, clear all AWArticleStore from sections and metadata (T429873), AWStorage: Use global stash keys (T430060) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 08:58 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 08:58 jforrester@deploy1003: Started scap sync-world: Backport for On AW article deletion, clear all AWArticleStore from sections and metadata (T429873), AWStorage: Use global stash keys (T430060)
- 08:57 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 08:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 08:55 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2234.codfw.wmnet with OS trixie
- 08:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2236.codfw.wmnet with OS trixie
- 08:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2236: Upgrading db2236.codfw.wmnet
- 08:52 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2236: Upgrading db2236.codfw.wmnet
- 08:52 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 08:52 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1221.eqiad.wmnet with OS trixie
- 08:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1221: Upgrading db1221.eqiad.wmnet
- 08:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1221: Upgrading db1221.eqiad.wmnet
- 08:47 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 08:47 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging2006.codfw.wmnet with OS trixie
- 08:45 marostegui@cumin1003: conftool action : set/weight=30; selector: name=clouddb1026.eqiad.wmnet
- 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1024-1025].eqiad.wmnet,db1155.eqiad.wmnet with reason: Reimaging db1221
- 08:10 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ec879e3] (releasing): T430110 deploy to Jenkins primary (duration: 00m 52s)
- 08:10 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ec879e3] (releasing): T430110 deploy to Jenkins primary
- 08:07 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ec879e3] (releasing): T430110 retry Jenkins secondary (duration: 00m 53s)
- 08:07 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ec879e3] (releasing): T430110 retry Jenkins secondary
- 08:03 marostegui: Pool clouddb1026:s1 with a bit of weight T409557
- 08:03 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1026.eqiad.wmnet,service=s1
- 08:02 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1026.eqiad.wmnet
- 07:52 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 07:52 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Allocate IPs for cloudvirt1077 - filippo@cumin1003"
- 07:52 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Allocate IPs for cloudvirt1077 - filippo@cumin1003"
- 07:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on releases2003.codfw.wmnet with reason: T410849
- 07:47 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 07:41 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db2160.codfw.wmnet with reason: Upgrading
- 07:35 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2234.codfw.wmnet with OS trixie
- 07:35 marostegui@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2234.codfw.wmnet with OS trixie
- 07:29 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@86ab691] (releasing): T430110 Test on Jenkins secondary (duration: 00m 50s)
- 07:29 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@86ab691] (releasing): T430110 Test on Jenkins secondary
- 07:24 moritzm: installing nginx security updates
- 07:21 dcausse: T423993: dropping ttmserver indices from the cirrussearch opensearch clusters
- 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2160.codfw.wmnet with reason: Upgrading
- 07:11 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2234.codfw.wmnet with OS trixie
- 07:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2234.codfw.wmnet,db1250.eqiad.wmnet with reason: Upgrading
- 07:07 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Cloning
- 07:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1026.eqiad.wmnet with reason: Catching up
- 06:40 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
- 06:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
- 06:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
- 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Upgrade gitlab
- 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
- 05:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on db1228.eqiad.wmnet with reason: Reimage to Trixie
- 05:46 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Upgrade gitlab
- 05:35 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 05:35 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 05:17 marostegui: Failover m2 from db1228 to db1290 - T429929
- 05:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2233].codfw.wmnet,db[1217,1228,1290].eqiad.wmnet with reason: Primary switchover m2 T429929
- 02:44 pt1979@cumin1003: END (PASS) - Cookbook sre.network.cf (exit_code=0)
- 02:44 pt1979@cumin1003: START - Cookbook sre.network.cf
- 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 24s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-24
- 23:28 tgr_: emergency deploy done
- 23:26 tgr@deploy1003: Finished scap sync-world: Backport for Fix B/C break for OAuth 1 format=json option, Fix B/C break for OAuth 1 format=json option, round #2 (duration: 12m 08s)
- 23:22 tgr@deploy1003: tgr: Continuing with deployment
- 23:16 tgr@deploy1003: tgr: Backport for Fix B/C break for OAuth 1 format=json option, Fix B/C break for OAuth 1 format=json option, round #2 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 23:14 tgr@deploy1003: Started scap sync-world: Backport for Fix B/C break for OAuth 1 format=json option, Fix B/C break for OAuth 1 format=json option, round #2
- 23:10 tgr_: doing an emergency backport for T430092
- 23:06 jdlrobson@deploy1003: Finished scap sync-world: Backport for Restore menu tab underline style (T428519), Reduce watchstar icon size (T426131), Replace Tools button with vertical ellipsis (T429258) (duration: 07m 21s)
- 23:02 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 23:01 jdlrobson@deploy1003: jdlrobson: Backport for Restore menu tab underline style (T428519), Reduce watchstar icon size (T426131), Replace Tools button with vertical ellipsis (T429258) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:59 jdlrobson@deploy1003: Started scap sync-world: Backport for Restore menu tab underline style (T428519), Reduce watchstar icon size (T426131), Replace Tools button with vertical ellipsis (T429258)
- 21:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:03 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Drop $wmgEmergencyCaptcha (T429849) (duration: 06m 43s)
- 20:59 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 20:58 dreamyjazz@deploy1003: dreamyjazz: Backport for Drop $wmgEmergencyCaptcha (T429849) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:56 dreamyjazz@deploy1003: Started scap sync-world: Backport for Drop $wmgEmergencyCaptcha (T429849)
- 20:37 XioNoX: bouncing cr1-eqiad FPC1 PIC1 - T429623
- 20:34 XioNoX: draining one of eqiad-codfw transports for PIC bounce
- 20:33 cscott@deploy1003: Finished scap sync-world: Backport for Add $wgParserMigrationEnableParsoid as unified/fine-grained config (duration: 08m 21s)
- 20:29 cscott@deploy1003: cscott: Continuing with deployment
- 20:28 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:28 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:27 XioNoX: cr1-eqiad# set chassis fpc 1 pic 1 port 5 speed 100g - T429623
- 20:27 cscott@deploy1003: cscott: Backport for Add $wgParserMigrationEnableParsoid as unified/fine-grained config synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:25 cscott@deploy1003: Started scap sync-world: Backport for Add $wgParserMigrationEnableParsoid as unified/fine-grained config
- 20:20 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 20:20 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 20:20 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 20:19 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 20:18 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:18 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:04 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:04 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 20:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 20:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Migration of db2219.codfw.wmnet completed
- 20:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 20:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Migration of db1242.eqiad.wmnet completed
- 19:20 jgreen@dns1004: END - running authdns-update
- 19:18 jgreen@dns1004: START - running authdns-update
- 19:18 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Migration of db2219.codfw.wmnet completed
- 19:15 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
- 19:15 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Migration of db1242.eqiad.wmnet completed
- 19:14 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
- 19:14 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
- 19:13 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
- 19:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2219.codfw.wmnet with OS trixie
- 19:02 swfrench-wmf: applied latent admin_ng diffs for allow-urldownloaders GlobalNetworkPolicy - T430045 T427282
- 19:02 swfrench-wmf: applied latent admin_ng diffs for mw-pretrain - T427668
- 19:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1242.eqiad.wmnet with OS trixie
- 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
- 18:56 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
- 18:49 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 18:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2219.codfw.wmnet with reason: host reimage
- 18:46 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 18:45 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 18:44 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1242.eqiad.wmnet with reason: host reimage
- 18:44 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2219.codfw.wmnet with reason: host reimage
- 18:43 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 18:43 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 18:40 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 18:40 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1242.eqiad.wmnet with reason: host reimage
- 18:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet with OS bookworm
- 18:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1015.eqiad.wmnet with OS bookworm
- 18:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet with OS bookworm
- 18:25 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2219.codfw.wmnet with OS trixie
- 18:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2219: Upgrading db2219.codfw.wmnet
- 18:23 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2219: Upgrading db2219.codfw.wmnet
- 18:23 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 18:23 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 18:22 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1242.eqiad.wmnet with OS trixie
- 18:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Upgrading db1242.eqiad.wmnet
- 18:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1016.eqiad.wmnet with reason: host reimage
- 18:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Upgrading db1242.eqiad.wmnet
- 18:21 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 18:21 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 18:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 18:19 kamila@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 18:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1015.eqiad.wmnet with reason: host reimage
- 18:16 kamila@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 18:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1017.eqiad.wmnet with reason: host reimage
- 18:12 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.8 refs T423917
- 18:10 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1016.eqiad.wmnet with reason: host reimage
- 18:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1015.eqiad.wmnet with reason: host reimage
- 18:08 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1017.eqiad.wmnet with reason: host reimage
- 18:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Upgrading db1242.eqiad.wmnet
- 18:05 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Upgrading db1242.eqiad.wmnet
- 18:05 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 18:05 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 18:01 brennen: 1.47.0-wmf.8 train status (T423917): no current blockers, logs no worse than expected, rolling to group1
- 17:55 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 17:55 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 17:54 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 17:54 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 17:52 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1017.eqiad.wmnet with OS bookworm
- 17:52 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1016.eqiad.wmnet with OS bookworm
- 17:52 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1015.eqiad.wmnet with OS bookworm
- 17:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T430072
- 17:46 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 17:46 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1016.eqiad.wmnet
- 17:46 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
- 17:45 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 17:45 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 17:45 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 17:42 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T430072
- 17:38 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 17:38 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T430072
- 17:36 kamila@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Migration of db2210.codfw.wmnet completed
- 17:35 kamila@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 17:35 kamila@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 17:34 kamila@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 17:34 kamila@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 17:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T430072
- 17:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Migration of db1241.eqiad.wmnet completed
- 17:06 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 17:05 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 16:51 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Migration of db2210.codfw.wmnet completed
- 16:41 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 16:41 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 16:37 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 16:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2210.codfw.wmnet with OS trixie
- 16:37 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 16:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Migration of db1241.eqiad.wmnet completed
- 16:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1241.eqiad.wmnet with OS trixie
- 16:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2210.codfw.wmnet with reason: host reimage
- 16:12 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2210.codfw.wmnet with reason: host reimage
- 16:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1241.eqiad.wmnet with reason: host reimage
- 15:58 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1241.eqiad.wmnet with reason: host reimage
- 15:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:54 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
- 15:53 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
- 15:53 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
- 15:52 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2210.codfw.wmnet with OS trixie
- 15:52 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
- 15:52 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:52 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:52 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
- 15:49 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:47 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
- 15:46 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
- 15:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1241.eqiad.wmnet with OS trixie
- 15:43 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2210: Upgrading db2210.codfw.wmnet
- 15:41 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
- 15:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Upgrading db1241.eqiad.wmnet
- 15:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2210: Upgrading db2210.codfw.wmnet
- 15:40 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 15:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:40 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Upgrading db1241.eqiad.wmnet
- 15:38 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 15:38 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:33 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
- 15:32 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
- 15:32 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
- 15:31 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
- 15:31 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:30 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:30 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
- 15:30 cgoubert@deploy1003: Finished scap sync-world: Backport for Update interwiki map (T429372 T418494) (duration: 14m 09s)
- 15:29 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
- 15:29 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
- 15:25 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
- 15:25 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
- 15:21 cgoubert@deploy1003: cgoubert: Continuing with deployment
- 15:20 cgoubert@deploy1003: cgoubert: Backport for Update interwiki map (T429372 T418494) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:16 cgoubert@deploy1003: Started scap sync-world: Backport for Update interwiki map (T429372 T418494)
- 15:15 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
- 15:14 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
- 15:12 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
- 15:11 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
- 15:11 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
- 15:11 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:11 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 15:10 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
- 15:10 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
- 15:10 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
- 15:10 mfossati@deploy1003: Finished scap sync-world: Backport for Restore the per-reader opt-out for the mobile image carousel (T419786), Restore the per-reader opt-out for the mobile image carousel (T419786) (duration: 35m 52s)
- 15:10 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
- 15:09 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
- 15:09 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
- 14:57 mfossati@deploy1003: egardner, mfossati: Continuing with deployment
- 14:53 mfossati@deploy1003: egardner, mfossati: Backport for Restore the per-reader opt-out for the mobile image carousel (T419786), Restore the per-reader opt-out for the mobile image carousel (T419786) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Migration of db2206.codfw.wmnet completed
- 14:47 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:47 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Migration of db1238.eqiad.wmnet completed
- 14:46 moritzm: installing postgresql security updates
- 14:34 mfossati@deploy1003: Started scap sync-world: Backport for Restore the per-reader opt-out for the mobile image carousel (T419786), Restore the per-reader opt-out for the mobile image carousel (T419786)
- 14:26 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 14:26 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp3074.esams.wmnet
- 14:26 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp3074.esams.wmnet
- 14:26 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp3066.esams.wmnet
- 14:26 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp3066.esams.wmnet
- 14:26 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 14:26 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.*
- 14:26 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:26 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp3066.*
- 14:26 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp3066.*
- 14:25 fabfur: repooling cp3066 and cp3074 after reimage (T419825)
- 14:25 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:25 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:25 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:24 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:24 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:17 Dreamy_Jazz: Afternoon UTC backport window done
- 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), Create ConfirmEditGetGlobalInstanceFromContext hook (T429848) (duration: 07m 48s)
- 14:16 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:16 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:16 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 14:15 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 14:15 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:14 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 14:11 dreamyjazz@deploy1003: dreamyjazz: Backport for Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), Create ConfirmEditGetGlobalInstanceFromContext hook (T429848) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:10 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 14:09 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:09 dreamyjazz@deploy1003: Started scap sync-world: Backport for Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), Create ConfirmEditGetGlobalInstanceFromContext hook (T429848)
- 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:07 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Migration of db2206.codfw.wmnet completed
- 14:05 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp3074.esams.wmnet
- 14:05 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp3074.esams.wmnet
- 14:05 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp3066.esams.wmnet
- 14:05 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp3066.esams.wmnet
- 14:01 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Migration of db1238.eqiad.wmnet completed
- {{safesubst:SAL entry|1=13:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Fix autonym for Khasi (kha) in wmgExtraLanguageNames (T427917), csbwiki: update logo, wordmark and tagline (T429126), Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), hCaptcha: Enable for Special:Contact (T429848), [[gerrit:1305408|Create ConfirmEditGetGlobalInstan}}
- 13:55 dreamyjazz@deploy1003: dreamyjazz, valn-ilyo, anzx: Continuing with deployment
- {{safesubst:SAL entry|1=13:52 dreamyjazz@deploy1003: dreamyjazz, valn-ilyo, anzx: Backport for Fix autonym for Khasi (kha) in wmgExtraLanguageNames (T427917), csbwiki: update logo, wordmark and tagline (T429126), Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), hCaptcha: Enable for Special:Contact (T429848), [[gerrit:1305408|Create ConfirmEditGetGlobalIns}}
- 13:51 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3066.esams.wmnet with OS trixie
- 13:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2206.codfw.wmnet with OS trixie
- {{safesubst:SAL entry|1=13:50 dreamyjazz@deploy1003: Started scap sync-world: Backport for Fix autonym for Khasi (kha) in wmgExtraLanguageNames (T427917), csbwiki: update logo, wordmark and tagline (T429126), Handle the ConfirmEditGetGlobalInstanceFromContext hook (T429848), hCaptcha: Enable for Special:Contact (T429848), [[gerrit:1305408|Create ConfirmEditGetGlobalInstanc}}
- 13:46 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
- 13:45 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1238.eqiad.wmnet with OS trixie
- 13:37 jgiannelos@deploy1003: Finished scap sync-world: Backport for Disable parser survey for all wikis (duration: 16m 01s)
- 13:33 jgiannelos@deploy1003: mbsantos, jgiannelos: Continuing with deployment
- 13:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2206.codfw.wmnet with reason: host reimage
- 13:32 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 13:29 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 13:28 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1238.eqiad.wmnet with reason: host reimage
- 13:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: repool after rack maintenance
- 13:26 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2206.codfw.wmnet with reason: host reimage
- 13:24 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3066.esams.wmnet with reason: host reimage
- 13:23 jgiannelos@deploy1003: mbsantos, jgiannelos: Backport for Disable parser survey for all wikis synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:22 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: rack depool
- 13:21 jgiannelos@deploy1003: Started scap sync-world: Backport for Disable parser survey for all wikis
- 13:19 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
- 13:15 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1238.eqiad.wmnet with reason: host reimage
- 13:15 kharlan@deploy1003: Finished scap sync-world: Backport for CheckUserGetUsersPager: Fix TypeError for numeric usernames (T429971) (duration: 09m 17s)
- 13:14 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3066.esams.wmnet with reason: host reimage
- 13:13 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
- 13:11 kharlan@deploy1003: kharlan: Continuing with deployment
- 13:08 kharlan@deploy1003: kharlan: Backport for CheckUserGetUsersPager: Fix TypeError for numeric usernames (T429971) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2206.codfw.wmnet with OS trixie
- 13:06 kharlan@deploy1003: Started scap sync-world: Backport for CheckUserGetUsersPager: Fix TypeError for numeric usernames (T429971)
- 13:01 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 13:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1238.eqiad.wmnet with OS trixie
- 12:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 12:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 12:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 12:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 9 hosts with reason: maintenance
- 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 12:51 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:51 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new VIP for dumps-nfs - filippo@cumin1003"
- 12:51 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new VIP for dumps-nfs - filippo@cumin1003"
- 12:49 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
- 12:49 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp3066.esams.wmnet with OS trixie
- 12:48 cwilliams@dns1005: END - running authdns-update
- 12:48 claime: Deleting apiportalwiki references in GlobalUsage - T418494
- 12:47 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 12:47 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.*
- 12:47 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp3066.*
- 12:47 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp3066.*
- 12:47 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 12:46 claime: Setting globaluser gu_home_db to NULL for apiportalwiki globalusers - T418494
- 12:46 fabfur: depooling cp3066 and cp3074 to reimage (T419825)
- 12:46 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:46 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new VIP for dumps-nfs - filippo@cumin1003"
- 12:46 cwilliams@dns1005: START - running authdns-update
- 12:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: new VIP for dumps-nfs - filippo@cumin1003"
- 12:45 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp2044.*
- 12:45 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp2043.*
- 12:44 fabfur: repooling cp2043 and cp2044 after reimage (T419825)
- 12:44 claime: Deleting apiportalwiki references in localnames table - T418494
- 12:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2155: repool after rack maintenance
- 12:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
- 12:40 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp2044.codfw.wmnet
- 12:40 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp2044.codfw.wmnet
- 12:40 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp2043.codfw.wmnet
- 12:40 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp2043.codfw.wmnet
- 12:39 ayounsi@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2155: rack depool
- 12:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2206: Upgrading db2206.codfw.wmnet
- 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2206: Upgrading db2206.codfw.wmnet
- 12:38 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 12:38 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1238: Upgrading db1238.eqiad.wmnet
- 12:38 claime: Deleting apiportalwiki references in localuser table - T418494
- 12:38 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1238: Upgrading db1238.eqiad.wmnet
- 12:37 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 12:37 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:36 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2156: rack depool
- 12:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2067,2072-2073,2114-2115,2124-2127,2256-2257].codfw.wmnet
- 12:35 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2067,2072-2073,2114-2115,2124-2127,2256-2257].codfw.wmnet
- 12:35 cgoubert@deploy1003: Finished scap sync-world: Backport for Remove config related to the API Portal (T429372 T418494), CommonSettings-labs: Remove api.wikimedia.beta.wmcloud.org (T429372 T418494) (duration: 08m 34s)
- 12:35 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2155: rack depool
- 12:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet
- 12:34 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet
- 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet
- 12:33 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet
- 12:31 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host db1208.eqiad.wmnet
- 12:31 cgoubert@deploy1003: apaskulin, cgoubert: Continuing with deployment
- 12:29 cgoubert@deploy1003: apaskulin, cgoubert: Backport for Remove config related to the API Portal (T429372 T418494), CommonSettings-labs: Remove api.wikimedia.beta.wmcloud.org (T429372 T418494) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:26 cgoubert@deploy1003: Started scap sync-world: Backport for Remove config related to the API Portal (T429372 T418494), CommonSettings-labs: Remove api.wikimedia.beta.wmcloud.org (T429372 T418494)
- 12:23 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Testing apiportalwiki deletion in beta - T418494 (duration: 98m 50s)
- 12:17 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet
- 12:17 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2067,2072-2073,2114-2115,2124-2127,2256-2257].codfw.wmnet
- 12:16 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
- 12:11 jmm@dns1004: END - running authdns-update
- 12:10 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 12:09 jmm@dns1004: START - running authdns-update
- 12:09 jmm@dns1004: END - running authdns-update
- 12:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet
- 12:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet
- 12:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet
- 12:07 jmm@dns1004: START - running authdns-update
- 12:06 dcausse: T423993: closing ttmserver indices in the cirrussearch opensearch cluster (eqiad & codfw)
- 12:05 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2067,2072-2073,2114-2115,2124-2127,2256-2257].codfw.wmnet
- 12:05 jmm@dns1004: END - running authdns-update
- 12:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2156: rack depool
- 12:04 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2156: rack depool
- 12:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2155: rack depool
- 12:03 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2155: rack depool
- 12:03 jmm@dns1004: START - running authdns-update
- 12:02 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for codfw rack A6
- 12:02 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-a6-codfw,lsw1-a6-codfw IPv6,lsw1-a6-codfw.mgmt with reason: Switch maintenance
- 12:01 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Switch maintenance
- 11:53 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack A6
- 11:53 moritzm: installing postgresql security updates
- 11:50 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1026.eqiad.wmnet
- 11:49 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1026.eqiad.wmnet,service=s1
- 11:48 jmm@dns1004: END - running authdns-update
- 11:48 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet
- 11:47 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
- 11:47 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
- 11:46 jmm@dns1004: START - running authdns-update
- 11:44 jmm@dns1004: END - running authdns-update
- 11:42 jmm@dns1004: START - running authdns-update
- 11:36 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 11:34 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 11:33 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 11:27 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 11:11 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 11:11 jmm@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 10:57 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2043.codfw.wmnet with OS trixie
- 10:51 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
- 10:45 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Testing apiportalwiki deletion in beta - T418494
- 10:34 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2043.codfw.wmnet with reason: host reimage
- 10:30 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
- 10:26 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2043.codfw.wmnet with reason: host reimage
- 10:26 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
- 10:10 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp2043.codfw.wmnet with OS trixie
- 10:10 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
- 10:05 fabfur: depooling cp2043 and cp2044 to reimage (T419825)
- 10:05 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp2044.*
- 10:05 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp2043.*
- 09:58 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1013.eqiad.wmnet,service=s1
- 09:54 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 09:54 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 09:45 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 09:45 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2179: Migration of db2179.codfw.wmnet completed
- {{safesubst:SAL entry|1=09:42 jforrester@deploy1003: Finished scap sync-world: Backport for [testwiki] Enable Abstract Client integration mode, not just previews (T422657), [abstractwiki] Add the 'allowed' temporary vars for cross-wiki content (T422657), WikiLambda: Expose wikilambda-abstract-optin for global group assignment (T422698), [[gerrit:1304110|[abstractwiki] Update favicon with new}}
- 09:37 jforrester@deploy1003: jforrester: Continuing with deployment
- {{safesubst:SAL entry|1=09:36 jforrester@deploy1003: jforrester: Backport for [testwiki] Enable Abstract Client integration mode, not just previews (T422657), [abstractwiki] Add the 'allowed' temporary vars for cross-wiki content (T422657), WikiLambda: Expose wikilambda-abstract-optin for global group assignment (T422698), [[gerrit:1304110|[abstractwiki] Update favicon with new version (T429}}
- {{safesubst:SAL entry|1=09:34 jforrester@deploy1003: Started scap sync-world: Backport for [testwiki] Enable Abstract Client integration mode, not just previews (T422657), [abstractwiki] Add the 'allowed' temporary vars for cross-wiki content (T422657), WikiLambda: Expose wikilambda-abstract-optin for global group assignment (T422698), [[gerrit:1304110|[abstractwiki] Update favicon with new}}
- 09:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 09:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Migration of db1199.eqiad.wmnet completed
- 09:27 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp7009.magru.wmnet
- 09:27 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp7009.magru.wmnet
- 09:27 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp7001.magru.wmnet
- 09:27 fabfur@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp7001.magru.wmnet
- 09:27 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp7001.*
- 09:27 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp7009.*
- 09:26 fabfur: repooling cp7001 and cp7009 after reimage (T419825)
- 09:25 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 09:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2046.codfw.wmnet
- 09:22 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet
- 09:18 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
- 09:08 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7001.magru.wmnet with OS trixie
- 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 09:03 moritzm: temporarily remove ganeti2028 from the codfw cluster T429817
- 08:59 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2179: Migration of db2179.codfw.wmnet completed
- 08:53 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
- 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet
- 08:50 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
- 08:46 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Migration of db1199.eqiad.wmnet completed
- 08:43 fabfur@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7001.magru.wmnet with reason: host reimage
- 08:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2179.codfw.wmnet with OS trixie
- 08:38 fabfur@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7001.magru.wmnet with reason: host reimage
- 08:33 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 08:32 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 08:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1199.eqiad.wmnet with OS trixie
- 08:27 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
- 08:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2179.codfw.wmnet with reason: host reimage
- 08:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1013.eqiad.wmnet with reason: Cloning cloddb1026
- 08:18 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1
- 08:16 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2179.codfw.wmnet with reason: host reimage
- 08:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1199.eqiad.wmnet with reason: host reimage
- 08:13 fabfur@cumin1003: START - Cookbook sre.hosts.reimage for host cp7001.magru.wmnet with OS trixie
- 08:10 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1199.eqiad.wmnet with reason: host reimage
- 08:09 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp7009.*
- 08:08 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp7001.*
- 08:08 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp7001.*
- 08:07 fabfur: depooling cp7001 and cp7009 to reimage (T419825)
- 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:58 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2179.codfw.wmnet with OS trixie
- 07:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1199.eqiad.wmnet with OS trixie
- 07:55 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2179: Upgrading db2179.codfw.wmnet
- 07:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:55 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2179: Upgrading db2179.codfw.wmnet
- 07:54 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 07:54 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 07:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Upgrading db1199.eqiad.wmnet
- 07:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:51 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Upgrading db1199.eqiad.wmnet
- 07:50 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 07:50 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 07:42 jmm@dns1004: END - running authdns-update
- 07:41 jmm@dns1004: START - running authdns-update
- 07:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 07:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: Migration of es1035.eqiad.wmnet completed
- 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet
- 07:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet
- 07:09 mlitn@deploy1003: Finished scap sync-world: Backport for Enable MMV carousel on enwiki (T429509) (duration: 07m 49s)
- 07:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet
- 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.changedisk (exit_code=0) for changing disk type of ml-staging-etcd2001.codfw.wmnet to drbd
- 07:04 mlitn@deploy1003: mlitn: Continuing with deployment
- 07:03 mlitn@deploy1003: mlitn: Backport for Enable MMV carousel on enwiki (T429509) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:01 mlitn@deploy1003: Started scap sync-world: Backport for Enable MMV carousel on enwiki (T429509)
- 06:57 jmm@cumin2003: START - Cookbook sre.ganeti.changedisk for changing disk type of ml-staging-etcd2001.codfw.wmnet to drbd
- 06:54 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet
- 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet
- 06:54 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2028.codfw.wmnet
- 06:54 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: Migration of es1035.eqiad.wmnet completed
- 06:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet
- 06:51 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet
- 06:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1035.eqiad.wmnet with OS trixie
- 06:45 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet
- 06:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1035.eqiad.wmnet with reason: host reimage
- 06:24 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1035.eqiad.wmnet with reason: host reimage
- 06:08 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1035.eqiad.wmnet with OS trixie
- 05:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: Upgrading es1035.eqiad.wmnet
- 05:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: Upgrading es1035.eqiad.wmnet
- 05:55 marostegui@cumin1003: dbmaint on es7@eqiad T429463
- 05:55 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 05:47 marostegui@dns1004: END - running authdns-update
- 05:46 marostegui@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - T429867', diff saved to https://phabricator.wikimedia.org/P94396 and previous config saved to /var/cache/conftool/dbconfig/20260624-054611-marostegui.json
- 05:45 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es1035 T429867', diff saved to https://phabricator.wikimedia.org/P94395 and previous config saved to /var/cache/conftool/dbconfig/20260624-054547-marostegui.json
- 05:45 marostegui@dns1004: START - running authdns-update
- 05:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote es1039 to es7 primary T429867', diff saved to https://phabricator.wikimedia.org/P94394 and previous config saved to /var/cache/conftool/dbconfig/20260624-054446-marostegui.json
- 05:44 marostegui: Starting es7 eqiad failover from es1035 to es1039 - T429867
- 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Set es1039 with weight 0 T429867', diff saved to https://phabricator.wikimedia.org/P94393 and previous config saved to /var/cache/conftool/dbconfig/20260624-054131-marostegui.json
- 05:41 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 T429867
- 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - T429867', diff saved to https://phabricator.wikimedia.org/P94392 and previous config saved to /var/cache/conftool/dbconfig/20260624-054106-marostegui.json
- 05:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis isvwiki in section s5
- 05:13 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis isvwiki in section s5
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-23
- 22:21 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 12389
- 22:20 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 12389
- 22:15 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Temporarily disable risk score block collection (T428659), hCaptcha: Reenable risk score block collection (T428659), Update HCaptchaRiskScoreRetrievedForBlocks hook to not use IDs (T428659) (duration: 09m 14s)
- 22:11 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 22:08 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Temporarily disable risk score block collection (T428659), hCaptcha: Reenable risk score block collection (T428659), Update HCaptchaRiskScoreRetrievedForBlocks hook to not use IDs (T428659) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:06 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Temporarily disable risk score block collection (T428659), hCaptcha: Reenable risk score block collection (T428659), Update HCaptchaRiskScoreRetrievedForBlocks hook to not use IDs (T428659)
- 21:23 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Revert "SourceEditorOverlay: Show CAPTCHA if has content in onStageChanges" (T429996) (duration: 06m 50s)
- 21:19 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 21:18 dreamyjazz@deploy1003: dreamyjazz: Backport for Revert "SourceEditorOverlay: Show CAPTCHA if has content in onStageChanges" (T429996) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:16 dreamyjazz@deploy1003: Started scap sync-world: Backport for Revert "SourceEditorOverlay: Show CAPTCHA if has content in onStageChanges" (T429996)
- 20:44 Dreamy_Jazz: Evening UTC backport window done
- 20:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Define login interface name for secureEnclave.js (T429963), hCaptcha: Define login interface name for secureEnclave.js (T429963), Improve click intent event logging and exposure tracking, Expand strip markers when they are present in attribute values (T383004) (duration: 0
- 20:37 dreamyjazz@deploy1003: dreamyjazz, wmde-fisch, arlolra: Continuing with deployment
- 20:35 dreamyjazz@deploy1003: dreamyjazz, wmde-fisch, arlolra: Backport for hCaptcha: Define login interface name for secureEnclave.js (T429963), hCaptcha: Define login interface name for secureEnclave.js (T429963), Improve click intent event logging and exposure tracking, Expand strip markers when they are present in attribute values (T383004) synce
- 20:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Define login interface name for secureEnclave.js (T429963), hCaptcha: Define login interface name for secureEnclave.js (T429963), Improve click intent event logging and exposure tracking, Expand strip markers when they are present in attribute values (T383004)
- 20:18 sbassett@deploy1003: Finished scap sync-world: Backport for Lazily reject pre-fix parser-cache entries for noreferrer/noopener links (T429090 T429244), Enable MMV carousel on non-en wikipedias (T429509) (duration: 10m 17s)
- 20:14 sbassett@deploy1003: sbassett, mlitn: Continuing with deployment
- 20:10 sbassett@deploy1003: sbassett, mlitn: Backport for Lazily reject pre-fix parser-cache entries for noreferrer/noopener links (T429090 T429244), Enable MMV carousel on non-en wikipedias (T429509) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:08 sbassett@deploy1003: Started scap sync-world: Backport for Lazily reject pre-fix parser-cache entries for noreferrer/noopener links (T429090 T429244), Enable MMV carousel on non-en wikipedias (T429509)
- 20:00 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 20:00 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 20:00 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 19:59 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 18:52 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.8 refs T423917
- 18:37 brennen@deploy1003: Finished scap sync-world: Backport for [parser] Rename mStripExtTags to useParsoidFragments (T429928) (duration: 10m 49s)
- 18:31 brennen@deploy1003: brennen, cscott: Continuing with deployment
- 18:28 brennen@deploy1003: brennen, cscott: Backport for [parser] Rename mStripExtTags to useParsoidFragments (T429928) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:26 brennen@deploy1003: Started scap sync-world: Backport for [parser] Rename mStripExtTags to useParsoidFragments (T429928)
- 18:03 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
- 18:02 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
- 18:01 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
- 18:01 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
- 18:00 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 18:00 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 17:59 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
- 17:59 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
- 17:59 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
- 17:58 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
- 17:57 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
- 17:56 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
- 17:48 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
- 17:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
- 17:46 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
- 17:46 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
- 17:45 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 17:45 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 17:44 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
- 17:44 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
- 17:44 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
- 17:43 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
- 17:42 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
- 17:41 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
- 17:41 reedy@deploy1003: Finished scap sync-world: Backport for Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965), Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965) (duration: 18m 03s)
- 17:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 17:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 17:36 reedy@deploy1003: reedy: Continuing with deployment
- 17:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 17:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 17:25 reedy@deploy1003: reedy: Backport for Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965), Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 17:23 reedy@deploy1003: Started scap sync-world: Backport for Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965), Updated guzzlehttp/guzzle from 7.12.1 to 7.12.3 (T429965), Upgrade guzzlehttp/* (T429965)
- 17:15 ladsgroup@deploy1003: Finished scap sync-world: Backport for Activate Wikipedia Interslavic (T429920) (duration: 06m 43s)
- 17:13 cmooney@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest1006.eqiad.wmnet with OS trixie
- 17:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 17:10 ladsgroup@deploy1003: ladsgroup: Backport for Activate Wikipedia Interslavic (T429920) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 17:08 ladsgroup@deploy1003: Started scap sync-world: Backport for Activate Wikipedia Interslavic (T429920)
- 17:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Migration of db2172.codfw.wmnet completed
- 17:01 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 17:00 cmooney@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest1006.eqiad.wmnet with OS trixie
- 17:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for Init Wikipedia Interslavic (T429920) (duration: 06m 55s)
- 16:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 16:55 ladsgroup@deploy1003: ladsgroup: Backport for Init Wikipedia Interslavic (T429920) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 16:53 ladsgroup@deploy1003: Started scap sync-world: Backport for Init Wikipedia Interslavic (T429920)
- 16:52 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
- 16:52 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
- 16:52 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
- 16:51 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
- 16:51 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
- 16:51 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
- 16:51 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
- 16:51 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
- 16:51 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
- 16:51 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
- 16:51 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
- 16:50 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
- 16:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Migration of db2172.codfw.wmnet completed
- 16:15 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 16:14 cmooney@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest1006.eqiad.wmnet with OS trixie
- 16:10 ladsgroup@dns1004: END - running authdns-update
- 16:08 ladsgroup@dns1004: START - running authdns-update
- 16:07 ladsgroup@dns1004: END - running authdns-update
- 16:07 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 16:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 16:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Migration of db1190.eqiad.wmnet completed
- 16:05 ladsgroup@dns1004: START - running authdns-update
- 16:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2172.codfw.wmnet with OS trixie
- 15:57 cmooney@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest1006.eqiad.wmnet with OS trixie
- 15:53 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for badlogin on all wikis (T429843) (duration: 07m 06s)
- 15:49 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 15:48 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for badlogin on all wikis (T429843) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for badlogin on all wikis (T429843)
- 15:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2172.codfw.wmnet with reason: host reimage
- 15:42 rzl: rzl@deploy1003:~$ kube-env mw-script-deploy eqiad; helm uninstall yhn94m3m # job completed successfully but helm release stuck in pending-install
- 15:38 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2172.codfw.wmnet with reason: host reimage
- 15:26 dreamyjazz@deploy1003: Finished scap sync-world: Backport for CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), ULS rewrite: Don't initialize IME and undo tooltip on Minerva skin (T429774) (duration: 10m 59s)
- 15:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Migration of db1190.eqiad.wmnet completed
- 15:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2172.codfw.wmnet with OS trixie
- 15:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 15:20 dreamyjazz@deploy1003: dreamyjazz, abi: Continuing with deployment
- 15:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 15:19 dreamyjazz@deploy1003: dreamyjazz, abi: Backport for CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), ULS rewrite: Don't initialize IME and undo tooltip on Minerva skin (T429774) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha
- 15:18 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 15:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2172: Upgrading db2172.codfw.wmnet
- 15:18 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 15:17 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2172: Upgrading db2172.codfw.wmnet
- 15:17 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 15:17 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 15:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 15:15 dreamyjazz@deploy1003: Started scap sync-world: Backport for CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), CaptchaFactory: Fallback config for badloginperuser from badlogin (T429902), ULS rewrite: Don't initialize IME and undo tooltip on Minerva skin (T429774)
- 15:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1190.eqiad.wmnet with OS trixie
- 15:09 kamila@deploy1003: Finished scap sync-world: switch default image to Debian Bookworm, disable Bullseye (duration: 29m 18s)
- 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@4d1f033]: deploy phab1004 for T429925 (duration: 00m 48s)
- 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@4d1f033]: deploy phab1004 for T429925
- 15:07 brennen@deploy1003: Finished deploy [phabricator/deployment@4d1f033]: deploy phab2003 for T429925 (duration: 00m 50s)
- 15:07 brennen@deploy1003: Started deploy [phabricator/deployment@4d1f033]: deploy phab2003 for T429925
- 15:06 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab1004.eqiad.wmnet with reason: deploy
- 15:06 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet with reason: deploy
- 14:56 SandraEbele_: Deployed Refinery as part of weekly deployment train
- 14:55 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1190.eqiad.wmnet with reason: host reimage
- 14:51 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1190.eqiad.wmnet with reason: host reimage
- 14:44 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:44 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Migration of db2155.codfw.wmnet completed
- 14:40 kamila@deploy1003: Started scap sync-world: switch default image to Debian Bookworm, disable Bullseye
- 14:37 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 14:36 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1190.eqiad.wmnet with OS trixie
- 14:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Upgrading db1190.eqiad.wmnet
- 14:34 cmooney@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest1006.eqiad.wmnet with OS trixie
- 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Upgrading db1190.eqiad.wmnet
- 14:34 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 14:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1160: Migration of db1160.eqiad.wmnet completed
- 14:30 ebysans@deploy1003: Finished deploy [analytics/refinery@83cc0ad] (thin): Regular analytics weekly train THIN [analytics/refinery@83cc0ad3] (duration: 01m 59s)
- 14:28 ebysans@deploy1003: Started deploy [analytics/refinery@83cc0ad] (thin): Regular analytics weekly train THIN [analytics/refinery@83cc0ad3]
- 14:26 ebysans@deploy1003: Finished deploy [analytics/refinery@83cc0ad]: Regular analytics weekly train [analytics/refinery@83cc0ad3] (duration: 05m 57s)
- 14:24 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 14:23 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 14:23 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 14:23 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 14:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 14:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 14:20 ebysans@deploy1003: Started deploy [analytics/refinery@83cc0ad]: Regular analytics weekly train [analytics/refinery@83cc0ad3]
- 14:20 ebysans@deploy1003: Finished deploy [analytics/refinery@83cc0ad] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@83cc0ad3] (duration: 02m 00s)
- 14:18 ebysans@deploy1003: Started deploy [analytics/refinery@83cc0ad] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@83cc0ad3]
- 14:10 atsuko@deploy1003: Finished scap sync-world: Backport for translate: remove CirrusSearch endpoints (T425377) (duration: 08m 59s)
- 14:05 atsuko@deploy1003: atsuko: Continuing with deployment
- 14:03 atsuko@deploy1003: atsuko: Backport for translate: remove CirrusSearch endpoints (T425377) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:01 atsuko@deploy1003: Started scap sync-world: Backport for translate: remove CirrusSearch endpoints (T425377)
- 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Migration of db2155.codfw.wmnet completed
- 13:58 cmooney@cumin1003: START - Cookbook sre.hosts.reimage for host sretest1006.eqiad.wmnet with OS trixie
- 13:53 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) sretest1006.mgmt.eqiad.wmnet on all recursors
- 13:53 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache sretest1006.mgmt.eqiad.wmnet on all recursors
- 13:52 cscott@deploy1003: Finished scap sync-world: Backport for [parser] Add configuration to return experimental PFragment types, [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624) (duration: 08m 29s)
- 13:52 brennen@deploy1003: Finished deploy [phabricator/deployment@a640ed9]: test deploy phab2003 (duration: 01m 22s)
- 13:50 brennen@deploy1003: Started deploy [phabricator/deployment@a640ed9]: test deploy phab2003
- 13:48 cscott@deploy1003: cscott: Continuing with deployment
- 13:47 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudcephosd1053
- 13:47 cmooney@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1053
- 13:46 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1160: Migration of db1160.eqiad.wmnet completed
- 13:46 cmooney@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1053
- 13:46 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudcephosd1053
- 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.cloud-host (exit_code=99) for host cloudcephosd10543
- 13:46 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudcephosd10543
- 13:46 cscott@deploy1003: cscott: Backport for [parser] Add configuration to return experimental PFragment types, [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be v
- 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudcephosd1054
- 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1054
- 13:45 cmooney@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
- 13:45 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudcephosd1054
- 13:44 cscott@deploy1003: Started scap sync-world: Backport for [parser] Add configuration to return experimental PFragment types, [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624)
- 13:42 cscott@deploy1003: Finished scap sync-world: Backport for [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624) (duration: 12m 31s)
- 13:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2155.codfw.wmnet with OS trixie
- 13:33 cscott@deploy1003: cscott: Continuing with deployment
- 13:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1160.eqiad.wmnet with OS trixie
- 13:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 13:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for sretest1006 - cmooney@cumin1003"
- 13:31 cscott@deploy1003: cscott: Backport for [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:29 cscott@deploy1003: Started scap sync-world: Backport for [parser] Return HeadingPFragments while preprocessing (T391624 T387520 T387521 T384490 T387374), [parser] Return ExtTagPFragments while preprocessing (T429624)
- 13:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for sretest1006 - cmooney@cumin1003"
- 13:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2155.codfw.wmnet with reason: host reimage
- 13:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 13:16 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2155.codfw.wmnet with reason: host reimage
- 13:16 jelto@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Test apt-mark hold
- 13:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1160.eqiad.wmnet with reason: host reimage
- 13:13 jelto@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Test apt-mark hold
- 13:07 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1160.eqiad.wmnet with reason: host reimage
- 12:58 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2155.codfw.wmnet with OS trixie
- 12:55 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2155: Upgrading db2155.codfw.wmnet
- 12:55 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2155: Upgrading db2155.codfw.wmnet
- 12:55 cwilliams@cumin1003: dbmaint on s4@codfw T429893
- 12:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:53 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1160.eqiad.wmnet with OS trixie
- 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1160: Upgrading db1160.eqiad.wmnet
- 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1160: Upgrading db1160.eqiad.wmnet
- 12:48 cwilliams@cumin1003: dbmaint on s4@eqiad T429893
- 12:48 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:48 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 12:48 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 12:46 sbisson@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' .
- 12:46 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
- 12:45 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
- 12:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
- 12:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
- 12:40 sbisson@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' .
- 12:39 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 12:39 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
- 12:36 sbisson@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' .
- 12:32 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 12:31 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
- 12:22 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-master
- 12:22 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging-ctrl2002.codfw.wmnet
- 12:22 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging-ctrl2002.codfw.wmnet
- 12:19 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging-ctrl2002.codfw.wmnet
- 12:18 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging-ctrl2002.codfw.wmnet
- 12:18 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging-ctrl2001.codfw.wmnet
- 12:18 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging-ctrl2001.codfw.wmnet
- 12:15 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging-ctrl2001.codfw.wmnet
- 12:15 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging-ctrl2001.codfw.wmnet
- 12:15 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-master
- 11:52 kharlan@deploy1003: Finished scap sync-world: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891) (duration: 06m 51s)
- 11:47 kharlan@deploy1003: kharlan: Continuing with deployment
- 11:47 kharlan@deploy1003: kharlan: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:45 kharlan@deploy1003: Started scap sync-world: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891)
- 11:36 kharlan@deploy1003: Finished scap sync-world: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891) (duration: 06m 37s)
- 11:32 kharlan@deploy1003: kharlan: Continuing with deployment
- 11:31 kharlan@deploy1003: kharlan: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:29 kharlan@deploy1003: Started scap sync-world: Backport for hCaptcha: Log sitekeys on a sitekey-mismatch error (T429891)
- 11:22 kharlan@deploy1003: Finished scap sync-world: Backport for Tally: truncate BLT names by character so the result stays encodable (T427104), Tally: truncate BLT names by character so the result stays encodable (T427104) (duration: 06m 48s)
- 11:18 kharlan@deploy1003: kharlan: Continuing with deployment
- 11:18 kharlan@deploy1003: kharlan: Backport for Tally: truncate BLT names by character so the result stays encodable (T427104), Tally: truncate BLT names by character so the result stays encodable (T427104) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:15 kharlan@deploy1003: Started scap sync-world: Backport for Tally: truncate BLT names by character so the result stays encodable (T427104), Tally: truncate BLT names by character so the result stays encodable (T427104)
- 10:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1217.eqiad.wmnet with reason: Cloning db1290
- 10:28 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 10:27 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
- 10:24 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 10:22 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 10:14 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 09:41 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1262: After HW issues T428832
- 09:21 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1384.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1383.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1382.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1381.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1380.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1379.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1378.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1377.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1376.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:13 blake@cumin1003: conftool action : set/pooled=yes; selector: name=wikikube-worker1375.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1384.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1383.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1382.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1381.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1380.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1379.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1378.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1377.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:08 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1376.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 09:07 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1201 (T426633)', diff saved to https://phabricator.wikimedia.org/P94359 and previous config saved to /var/cache/conftool/dbconfig/20260623-090752-fceratto.json
- 09:07 blake@cumin1003: conftool action : set/weight=10; selector: name=wikikube-worker1375.eqiad.wmnet,cluster=kubernetes,service=kubesvc
- 08:57 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1201', diff saved to https://phabricator.wikimedia.org/P94357 and previous config saved to /var/cache/conftool/dbconfig/20260623-085744-fceratto.json
- 08:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1262: After HW issues T428832
- 08:48 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1262: After HW issues T428832
- 08:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1262: After HW issues T428832
- 08:47 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1201', diff saved to https://phabricator.wikimedia.org/P94355 and previous config saved to /var/cache/conftool/dbconfig/20260623-084737-fceratto.json
- 08:37 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1201 (T426633)', diff saved to https://phabricator.wikimedia.org/P94354 and previous config saved to /var/cache/conftool/dbconfig/20260623-083729-fceratto.json
- 08:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: test
- 08:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: test
- 08:33 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1018: test
- 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: test
- 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Maintenance on pc8
- 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Maintenance on pc8
- 08:30 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db1201 (T426633)', diff saved to https://phabricator.wikimedia.org/P94351 and previous config saved to /var/cache/conftool/dbconfig/20260623-083046-fceratto.json
- 08:30 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance
- 07:46 Msz2001: Deployed update to SI private code
- 07:35 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: repool after recloning another host
- 07:35 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 07:34 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 07:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: repool after recloning another host
- 07:34 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1018: repool after recloning another host
- 07:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: repool after recloning another host
- 07:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Maintenance on pc2
- 07:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 07:34 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 07:34 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Maintenance on pc2
- 07:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: repool after recloning another host
- 07:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: repool after recloning another host
- 07:30 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc2018: repool after recloning another host
- 07:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2018: repool after recloning another host
- 07:30 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1018: repool after recloning another host
- 07:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: repool after recloning another host
- 07:23 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1018: after reimage to trixie
- 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: after reimage to trixie
- 07:22 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pc1018.eqiad.wmnet with OS trixie
- 07:15 wmde-fisch@deploy1003: Finished scap sync-world: Backport for Global rollout - Sub-ref deployments to group 2 wikis (batch 1) (T428902) (duration: 12m 37s)
- 07:08 wmde-fisch@deploy1003: lilients, wmde-fisch: Continuing with deployment
- 07:07 wmde-fisch@deploy1003: lilients, wmde-fisch: Backport for Global rollout - Sub-ref deployments to group 2 wikis (batch 1) (T428902) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:02 wmde-fisch@deploy1003: Started scap sync-world: Backport for Global rollout - Sub-ref deployments to group 2 wikis (batch 1) (T428902)
- 07:01 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Repooling after switchover
- 07:00 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pc1018.eqiad.wmnet with reason: host reimage
- 06:53 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on pc1018.eqiad.wmnet with reason: host reimage
- 06:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 06:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Migration of es2038.codfw.wmnet completed
- 06:37 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host pc1018.eqiad.wmnet with OS trixie
- 06:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Reimage to Trixie
- 06:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 06:32 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 06:32 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Reimage to Trixie
- 06:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on pc1018.eqiad.wmnet with reason: Reimage to Trixie
- 06:28 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc2018: after reimage to trixie
- 06:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2018: after reimage to trixie
- 06:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pc2018.codfw.wmnet with OS trixie
- 06:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Repooling after switchover
- 06:14 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db1201 T429868', diff saved to https://phabricator.wikimedia.org/P94338 and previous config saved to /var/cache/conftool/dbconfig/20260623-061416-fceratto.json
- 06:13 fceratto@dns1004: END - running authdns-update
- 06:11 fceratto@dns1004: START - running authdns-update
- 06:07 fceratto@cumin1003: dbctl commit (dc=all): 'Set s6 eqiad as read-only for maintenance - T429868', diff saved to https://phabricator.wikimedia.org/P94336 and previous config saved to /var/cache/conftool/dbconfig/20260623-060708-fceratto.json
- 06:06 federico3: Starting s6 eqiad failover from db1201 to db1173 - T429868
- 06:04 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pc2018.codfw.wmnet with reason: host reimage
- 06:01 fceratto@cumin1003: dbctl commit (dc=all): 'Set db1173 with weight 0 T429868', diff saved to https://phabricator.wikimedia.org/P94335 and previous config saved to /var/cache/conftool/dbconfig/20260623-060157-fceratto.json
- 06:01 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 T429868
- 06:00 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Migration of es2038.codfw.wmnet completed
- 05:59 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on pc2018.codfw.wmnet with reason: host reimage
- 05:53 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2038.codfw.wmnet with OS trixie
- 05:42 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host pc2018.codfw.wmnet with OS trixie
- 05:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2018: Reimage to Trixie
- 05:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 05:41 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 05:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2018: Reimage to Trixie
- 05:41 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on pc2018.codfw.wmnet with reason: Reimage to Trixie
- 05:35 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2038.codfw.wmnet with reason: host reimage
- 05:31 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2038.codfw.wmnet with reason: host reimage
- 05:13 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2038.codfw.wmnet with OS trixie
- 05:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2038: Upgrading es2038.codfw.wmnet
- 05:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2038: Upgrading es2038.codfw.wmnet
- 05:12 marostegui@cumin1003: dbmaint on es7@codfw T429463
- 05:12 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 05:10 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es2038 T429794', diff saved to https://phabricator.wikimedia.org/P94332 and previous config saved to /var/cache/conftool/dbconfig/20260623-051012-marostegui.json
- 05:08 marostegui@cumin1003: dbctl commit (dc=all): 'Promote es2039 to es7 primary T429794', diff saved to https://phabricator.wikimedia.org/P94331 and previous config saved to /var/cache/conftool/dbconfig/20260623-050758-marostegui.json
- 05:07 marostegui: Starting es7 codfw failover from es2038 to es2039 - T429794
- 05:01 marostegui@cumin1003: dbctl commit (dc=all): 'Set es2039 with weight 0 T429794', diff saved to https://phabricator.wikimedia.org/P94330 and previous config saved to /var/cache/conftool/dbconfig/20260623-050137-marostegui.json
- 05:01 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 T429794
- 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.5 (duration: 02m 37s)
- 03:44 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.8 refs T423917 (duration: 39m 23s)
- 03:08 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.8 refs T423917
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 25s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:41 egardner@deploy1003: Finished scap sync-world: Backport for Remove multimediaviewer-beta (duration: 11m 14s)
- 01:34 egardner@deploy1003: egardner, ksarabia: Continuing with deployment
- 01:34 egardner@deploy1003: egardner, ksarabia: Backport for Remove multimediaviewer-beta synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 01:30 egardner@deploy1003: Started scap sync-world: Backport for Remove multimediaviewer-beta
- 01:24 egardner@deploy1003: Finished scap sync-world: Backport for Inject service RepoGroup into Hooks, MMV Beta Viewer: Improve loading/navigation UX (T429193), Take the feature out of beta (T429509) (duration: 33m 22s)
- 01:12 egardner@deploy1003: egardner: Continuing with deployment
- 01:10 egardner@deploy1003: egardner: Backport for Inject service RepoGroup into Hooks, MMV Beta Viewer: Improve loading/navigation UX (T429193), Take the feature out of beta (T429509) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 01:10 brett@dns7002: END - running authdns-update
- 01:08 brett@dns7002: START - running authdns-update
- 00:57 brett@dns7002: END - running authdns-update
- 00:56 brett@dns7002: START - running authdns-update
- 00:51 egardner@deploy1003: Started scap sync-world: Backport for Inject service RepoGroup into Hooks, MMV Beta Viewer: Improve loading/navigation UX (T429193), Take the feature out of beta (T429509)
- 00:46 swfrench-wmf: manually started scap-clean-images.service on deply1003 to reclaim /srv space (93% -> 66% utilization)
- 00:39 mutante: codesearch10: systemctl start codesearch-write-config; systemctl restart hound-operations (gerrit:1304848) (T429819)
- 00:33 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1027.eqiad.wmnet with OS trixie
- 00:33 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:33 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:30 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1033.eqiad.wmnet with OS trixie
- 00:30 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:24 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1026.eqiad.wmnet with OS trixie
- 00:24 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:24 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:20 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1028.eqiad.wmnet with OS trixie
- 00:20 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:20 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:17 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1027.eqiad.wmnet with reason: host reimage
- 00:15 egardner@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.6,1.47.0-wmf.7,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/me
- 00:14 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1033.eqiad.wmnet with reason: host reimage
- 00:11 egardner@deploy1003: Started scap sync-world: Backport for Inject service RepoGroup into Hooks, MMV Beta Viewer: Improve loading/navigation UX (T429193), Take the feature out of beta (T429509)
- 00:10 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1029.eqiad.wmnet with OS trixie
- 00:10 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:10 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 00:09 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1026.eqiad.wmnet with reason: host reimage
- 00:08 egardner@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.6,1.47.0-wmf.7,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/me
- 00:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1027.eqiad.wmnet with reason: host reimage
- 00:05 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1028.eqiad.wmnet with reason: host reimage
- 00:03 egardner@deploy1003: Started scap sync-world: Backport for Inject service RepoGroup into Hooks, MMV Beta Viewer: Improve loading/navigation UX (T429193), Take the feature out of beta (T429509)
- 00:03 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1033.eqiad.wmnet with reason: host reimage
2026-06-22
- 23:57 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1026.eqiad.wmnet with reason: host reimage
- 23:57 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1027.eqiad.wmnet with OS trixie
- 23:57 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1028.eqiad.wmnet with reason: host reimage
- 23:55 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1029.eqiad.wmnet with reason: host reimage
- 23:53 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1033.eqiad.wmnet with OS trixie
- 23:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host clouddb1033.eqiad.wmnet with OS trixie
- 23:51 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1031.eqiad.wmnet with OS trixie
- 23:51 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:51 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1029.eqiad.wmnet with reason: host reimage
- 23:50 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:50 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:47 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1026.eqiad.wmnet with OS trixie
- 23:47 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1030.eqiad.wmnet with OS trixie
- 23:47 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:47 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:47 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1028.eqiad.wmnet with OS trixie
- 23:40 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1029.eqiad.wmnet with OS trixie
- 23:40 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddb1032.eqiad.wmnet with OS trixie
- 23:40 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:39 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
- 23:38 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:36 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1031.eqiad.wmnet with reason: host reimage
- 23:34 jdlrobson@deploy1003: Finished scap sync-world: Backport for Ensure page tools icons are only shown on small viewports (T426131), Fix duplicate print and other projects menu links in main menu (T429676) (duration: 08m 18s)
- 23:32 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1030.eqiad.wmnet with reason: host reimage
- 23:31 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:29 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 23:27 jdlrobson@deploy1003: jdlrobson: Backport for Ensure page tools icons are only shown on small viewports (T426131), Fix duplicate print and other projects menu links in main menu (T429676) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 23:27 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1031.eqiad.wmnet with reason: host reimage
- 23:27 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1030.eqiad.wmnet with reason: host reimage
- 23:26 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:25 jdlrobson@deploy1003: Started scap sync-world: Backport for Ensure page tools icons are only shown on small viewports (T426131), Fix duplicate print and other projects menu links in main menu (T429676)
- 23:25 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddb1032.eqiad.wmnet with reason: host reimage
- 23:21 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddb1032.eqiad.wmnet with reason: host reimage
- 23:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1030.eqiad.wmnet with OS trixie
- 23:05 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1031.eqiad.wmnet with OS trixie
- 23:05 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1033.eqiad.wmnet with OS trixie
- 23:04 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host clouddb1032.eqiad.wmnet with OS trixie
- 23:04 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1033.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:04 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1032.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:03 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1031.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 23:03 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1030.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:55 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1033.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:55 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1032.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:54 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1028.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1031.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:54 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1029.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:53 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1030.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:49 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host clouddb1026.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:48 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for bad login on group1 (T429843) (duration: 06m 45s)
- 22:44 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1029.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:43 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1028.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:41 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:41 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 22:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:40 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for bad login on group1 (T429843) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1027.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:40 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 22:39 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding clouddb1031 to eqiad - jclark@cumin1003"
- 22:39 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding clouddb1031 to eqiad - jclark@cumin1003"
- 22:39 jclark@cumin1003: START - Cookbook sre.hosts.provision for host clouddb1026.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 22:38 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for bad login on group1 (T429843)
- 22:35 jclark@cumin1003: START - Cookbook sre.dns.netbox
- 22:31 cdobbins@dns1004: END - running authdns-update
- 22:29 cdobbins@dns1004: START - running authdns-update
- 22:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Apply generic settings for bad login (T429843), hCaptcha: Enable for badlogin on non-SUL wikis (T429843) (duration: 07m 14s)
- 22:13 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 22:12 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Apply generic settings for bad login (T429843), hCaptcha: Enable for badlogin on non-SUL wikis (T429843) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Apply generic settings for bad login (T429843), hCaptcha: Enable for badlogin on non-SUL wikis (T429843)
- 22:04 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for badlogin on non-SUL wikis (T429843) (duration: 11m 37s)
- 22:04 dreamyjazz@deploy1003: dreamyjazz: Rolling back deployment
- 21:55 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for badlogin on non-SUL wikis (T429843) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:52 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for badlogin on non-SUL wikis (T429843)
- 21:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for RiskScoreCollector: Make error_context a string map (T429594), CaptchaPreAuthenticationProvider: Clear solved state on failure (T429705) (duration: 11m 12s)
- 21:36 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 21:36 dreamyjazz@deploy1003: dreamyjazz: Backport for RiskScoreCollector: Make error_context a string map (T429594), CaptchaPreAuthenticationProvider: Clear solved state on failure (T429705) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for RiskScoreCollector: Make error_context a string map (T429594), CaptchaPreAuthenticationProvider: Clear solved state on failure (T429705)
- 21:20 arlolra@deploy1003: Finished scap sync-world: Backport for Add a hidden lint for pre ext tags expanding templates (T353697) (duration: 33m 42s)
- 21:20 cdobbins@dns1004: END - running authdns-update
- 21:18 cdobbins@dns1004: START - running authdns-update
- 21:08 arlolra@deploy1003: arlolra: Continuing with deployment
- 21:06 arlolra@deploy1003: arlolra: Backport for Add a hidden lint for pre ext tags expanding templates (T353697) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:47 arlolra@deploy1003: Started scap sync-world: Backport for Add a hidden lint for pre ext tags expanding templates (T353697)
- 20:34 bpirkle@deploy1003: Finished scap sync-world: Backport for REST: adjust analytics and wikifunctions REST Sandbox visibility (T422770 T423058 T422771) (duration: 07m 30s)
- 20:30 bpirkle@deploy1003: bpirkle: Continuing with deployment
- 20:29 bpirkle@deploy1003: bpirkle: Backport for REST: adjust analytics and wikifunctions REST Sandbox visibility (T422770 T423058 T422771) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:27 bpirkle@deploy1003: Started scap sync-world: Backport for REST: adjust analytics and wikifunctions REST Sandbox visibility (T422770 T423058 T422771)
- 20:19 catrope@deploy1003: Finished scap sync-world: Backport for Permissions: Create wmf-officeit group on collabwiki (duration: 09m 45s)
- 20:15 catrope@deploy1003: catrope: Continuing with deployment
- 20:11 catrope@deploy1003: catrope: Backport for Permissions: Create wmf-officeit group on collabwiki synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:09 catrope@deploy1003: Started scap sync-world: Backport for Permissions: Create wmf-officeit group on collabwiki
- 19:34 reedy@deploy1003: Finished scap sync-world: Backport for Add info-level logging to wmgMonologChannels for timeline (T429654) (duration: 06m 41s)
- 19:29 reedy@deploy1003: sbassett, reedy: Continuing with deployment
- 19:29 reedy@deploy1003: sbassett, reedy: Backport for Add info-level logging to wmgMonologChannels for timeline (T429654) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 19:27 reedy@deploy1003: Started scap sync-world: Backport for Add info-level logging to wmgMonologChannels for timeline (T429654)
- 19:21 reedy@deploy1003: Finished scap sync-world: Backport for Upgrading guzzlehttp/psr7 (2.11.0 => 2.11.1), Upgrade guzzle/* (T429826), Updated guzzlehttp/guzzle from 7.10.0 to 7.12.1 (T429826) (duration: 09m 06s)
- 19:14 reedy@deploy1003: reedy: Continuing with deployment
- 19:14 reedy@deploy1003: reedy: Backport for Upgrading guzzlehttp/psr7 (2.11.0 => 2.11.1), Upgrade guzzle/* (T429826), Updated guzzlehttp/guzzle from 7.10.0 to 7.12.1 (T429826) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 19:12 reedy@deploy1003: Started scap sync-world: Backport for Upgrading guzzlehttp/psr7 (2.11.0 => 2.11.1), Upgrade guzzle/* (T429826), Updated guzzlehttp/guzzle from 7.10.0 to 7.12.1 (T429826)
- 16:47 kamila@deploy1003: Finished scap sync-world: upgrade all of prod to debian bookworm (duration: 09m 17s)
- 16:46 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 16:46 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 16:44 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 16:43 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 16:43 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 16:43 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 16:38 kamila@deploy1003: Started scap sync-world: upgrade all of prod to debian bookworm
- 16:37 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 16:31 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-k8s # T425377 populating translation memory (dblist: https://phabricator.wikimedia.org/P94329)
- 16:30 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 16:30 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 16:23 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 15:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1383.eqiad.wmnet with OS trixie
- 15:45 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1384.eqiad.wmnet with OS trixie
- 15:41 tgr_: UTC afternoon deploys double done
- 15:41 tgr@deploy1003: Finished scap sync-world: Backport for Fix displaying events with IP agents (T428198), Preserve redoLocalAuthentication flag when returning from auth domain (T429495) (duration: 14m 15s)
- 15:39 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1380.eqiad.wmnet with OS trixie
- 15:35 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1381.eqiad.wmnet with OS trixie
- 15:34 tgr@deploy1003: matmarex, tgr: Continuing with deployment
- 15:31 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1383.eqiad.wmnet with reason: host reimage
- 15:31 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1382.eqiad.wmnet with OS trixie
- 15:31 tgr@deploy1003: matmarex, tgr: Backport for Fix displaying events with IP agents (T428198), Preserve redoLocalAuthentication flag when returning from auth domain (T429495) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1384.eqiad.wmnet with reason: host reimage
- 15:27 tgr@deploy1003: Started scap sync-world: Backport for Fix displaying events with IP agents (T428198), Preserve redoLocalAuthentication flag when returning from auth domain (T429495)
- 15:22 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1380.eqiad.wmnet with reason: host reimage
- 15:17 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1381.eqiad.wmnet with reason: host reimage
- 15:13 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1382.eqiad.wmnet with reason: host reimage
- 15:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 15:11 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 15:09 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 15:09 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 15:09 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1383.eqiad.wmnet with reason: host reimage
- 15:09 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1384.eqiad.wmnet with reason: host reimage
- 15:08 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1380.eqiad.wmnet with reason: host reimage
- 15:08 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1381.eqiad.wmnet with reason: host reimage
- 15:08 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1382.eqiad.wmnet with reason: host reimage
- 15:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:04 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Migration of es1039.eqiad.wmnet completed
- 15:02 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 15:02 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 14:55 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1384.eqiad.wmnet with OS trixie
- 14:55 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1383.eqiad.wmnet with OS trixie
- 14:55 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1382.eqiad.wmnet with OS trixie
- 14:55 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1381.eqiad.wmnet with OS trixie
- 14:55 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1380.eqiad.wmnet with OS trixie
- 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1376.eqiad.wmnet with OS trixie
- 14:52 elukey: deny ANONYMOUS write traffic on Kafka main for statsv (only varnishkafka will push events from now on) - T425528
- 14:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1377.eqiad.wmnet with OS trixie
- 14:45 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1378.eqiad.wmnet with OS trixie
- 14:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1379.eqiad.wmnet with OS trixie
- 14:39 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1375.eqiad.wmnet with OS trixie
- 14:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1376.eqiad.wmnet with reason: host reimage
- 14:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1377.eqiad.wmnet with reason: host reimage
- 14:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1378.eqiad.wmnet with reason: host reimage
- 14:23 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1379.eqiad.wmnet with reason: host reimage
- 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1375.eqiad.wmnet with reason: host reimage
- 14:20 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1379.eqiad.wmnet with reason: host reimage
- 14:20 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1377.eqiad.wmnet with reason: host reimage
- 14:20 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1378.eqiad.wmnet with reason: host reimage
- 14:19 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1376.eqiad.wmnet with reason: host reimage
- 14:19 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Migration of es1039.eqiad.wmnet completed
- 14:16 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1375.eqiad.wmnet with reason: host reimage
- 14:08 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 14:07 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 14:07 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 14:07 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 14:07 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:07 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1379.eqiad.wmnet with OS trixie
- 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:07 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1378.eqiad.wmnet with OS trixie
- 14:06 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1039.eqiad.wmnet with OS trixie
- 14:06 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1377.eqiad.wmnet with OS trixie
- 14:06 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1376.eqiad.wmnet with OS trixie
- 14:04 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1375.eqiad.wmnet with OS trixie
- 14:03 tgr_: UTC afternoon deploys done(ish; will do the remaining two patches in about an hour)
- 14:02 tgr@deploy1003: Finished scap sync-world: Backport for config: Enable EmailConfirmationBanner on all wikis (T428292) (duration: 13m 08s)
- 13:56 tgr@deploy1003: mmartorana, tgr: Continuing with deployment
- 13:55 tgr@deploy1003: mmartorana, tgr: Backport for config: Enable EmailConfirmationBanner on all wikis (T428292) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 13:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1039.eqiad.wmnet with reason: host reimage
- 13:49 tgr@deploy1003: Started scap sync-world: Backport for config: Enable EmailConfirmationBanner on all wikis (T428292)
- 13:46 tgr@deploy1003: Finished scap sync-world: Backport for Add email confirmation banner Test Kitchen instrumentation (long-term) (T428293) (duration: 11m 59s)
- 13:43 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1039.eqiad.wmnet with reason: host reimage
- 13:40 tgr@deploy1003: tgr, mmartorana: Continuing with deployment
- 13:39 tgr@deploy1003: tgr, mmartorana: Backport for Add email confirmation banner Test Kitchen instrumentation (long-term) (T428293) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:34 tgr@deploy1003: Started scap sync-world: Backport for Add email confirmation banner Test Kitchen instrumentation (long-term) (T428293)
- 13:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 13:34 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 13:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 13:33 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 13:33 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 13:32 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 13:29 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 13:27 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
- 13:27 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1039.eqiad.wmnet with OS trixie
- 13:26 tchin@deploy1003: Finished scap sync-world: Backport for [PageViewInfo] Add new config (T411771) (duration: 14m 24s)
- 13:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1039: Upgrading es1039.eqiad.wmnet
- 13:25 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1039: Upgrading es1039.eqiad.wmnet
- 13:24 marostegui@cumin1003: dbmaint on es7@eqiad T429463
- 13:24 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 13:19 tchin@deploy1003: tchin: Continuing with deployment
- 13:17 tchin@deploy1003: tchin: Backport for [PageViewInfo] Add new config (T411771) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:11 tchin@deploy1003: Started scap sync-world: Backport for [PageViewInfo] Add new config (T411771)
- 13:04 oblivian@cumin2003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix utf-8 name handling - oblivian@cumin2003"
- 13:04 oblivian@cumin2003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix utf-8 name handling - oblivian@cumin2003
- 13:03 oblivian@cumin2003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix utf-8 name handling - oblivian@cumin2003
- 13:03 oblivian@cumin2003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix utf-8 name handling - oblivian@cumin2003"
- 13:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2212.codfw.wmnet
- 13:03 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2212.codfw.wmnet
- 13:02 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet
- 13:02 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet
- 12:55 Msz2001: Deployed changes to private code (SuggestedInvestigations and PrivateSettings.php)
- 12:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Migration of es1040.eqiad.wmnet completed
- 12:38 kharlan@deploy1003: Finished scap sync-world: Backport for hCaptcha: Skip blocked-IP risk-score collection for crawlers (T429755) (duration: 16m 20s)
- 12:33 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver codfw-k8s # T425377 populating translation memory (dblist: https://phabricator.wikimedia.org/P94319)
- 12:33 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-k8s # T425377 populating translation memory (dblist: https://phabricator.wikimedia.org/P94318)
- 12:32 kharlan@deploy1003: kharlan: Continuing with deployment
- 12:28 kharlan@deploy1003: kharlan: Backport for hCaptcha: Skip blocked-IP risk-score collection for crawlers (T429755) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:22 kharlan@deploy1003: Started scap sync-world: Backport for hCaptcha: Skip blocked-IP risk-score collection for crawlers (T429755)
- 12:19 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
- 12:18 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
- 12:15 kamila@deploy1003: Finished scap sync-world: upgrade canaries to Debian Bookworm (duration: 10m 33s)
- 12:14 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply
- 12:13 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply
- 12:12 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/ratelimit: apply
- 12:12 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/ratelimit: apply
- 12:06 kamila@deploy1003: Started scap sync-world: upgrade canaries to Debian Bookworm
- 11:53 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
- 11:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Migration of es1040.eqiad.wmnet completed
- 11:52 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
- 11:52 kamila@deploy1003: Finished scap sync-world: Upgrade mw-debug,mw-experimental to debian bookworm (duration: 33m 39s)
- 11:51 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply
- 11:50 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply
- 11:49 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
- 11:48 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
- 11:48 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply
- 11:48 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply
- 11:47 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/ratelimit: apply
- 11:47 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/ratelimit: apply
- 11:47 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/ratelimit: apply
- 11:47 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/ratelimit: apply
- 11:41 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1040.eqiad.wmnet with OS trixie
- 11:24 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1040.eqiad.wmnet with reason: host reimage
- 11:19 kamila@deploy1003: Started scap sync-world: Upgrade mw-debug,mw-experimental to debian bookworm
- 11:19 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1040.eqiad.wmnet with reason: host reimage
- 11:04 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1040.eqiad.wmnet with OS trixie
- 11:02 kamila@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.6,1.47.0-wmf.7,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/medi
- 11:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1040: Upgrading es1040.eqiad.wmnet
- 10:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1040: Upgrading es1040.eqiad.wmnet
- 10:59 marostegui@cumin1003: dbmaint on es7@eqiad T429463
- 10:59 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:47 kamila@deploy1003: Started scap sync-world: Upgrading mw-debug,mw-experimental to Debian Bookworm
- 10:45 Raine: upgrading mw-debug+mw-experimental to Bookworm
- 10:36 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 10:22 cgoubert@deploy1003: Finished scap sync-world: T418492 Redirect API Portal wiki URLs to www.mediawiki.org/wiki/Wikimedia_APIs (duration: 07m 17s)
- 10:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: sync
- 10:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: sync
- 10:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: sync
- 10:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: repool after recloning another host
- 10:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: sync
- 10:18 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: sync
- 10:18 cgoubert@deploy1003: cgoubert: Continuing with deployment
- 10:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: sync
- 10:16 cgoubert@deploy1003: cgoubert: T418492 Redirect API Portal wiki URLs to www.mediawiki.org/wiki/Wikimedia_APIs synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: repool after maintenance
- 10:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 10:16 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 10:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2017: repool after maintenance
- 10:15 cgoubert@deploy1003: Started scap sync-world: T418492 Redirect API Portal wiki URLs to www.mediawiki.org/wiki/Wikimedia_APIs
- 10:14 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc2017: repool after maintenance
- 10:14 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2017: repool after maintenance
- 10:14 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1017: repool after maintenance
- 10:14 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: repool after maintenance
- 10:13 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1017: repool after maintenance
- 10:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: repool after maintenance
- 09:58 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1017: after reimage to trixie
- 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: after reimage to trixie
- 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pc1017.eqiad.wmnet with OS trixie
- 09:52 jforrester@deploy1003: Finished scap sync-world: Backport for [abstractwiki] Enable Abstract Client mode (and on test wiki) (T422657) (duration: 08m 11s)
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Migration of es1048.eqiad.wmnet completed
- 09:48 jforrester@deploy1003: jforrester: Continuing with deployment
- 09:46 jforrester@deploy1003: jforrester: Backport for [abstractwiki] Enable Abstract Client mode (and on test wiki) (T422657) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:44 jforrester@deploy1003: Started scap sync-world: Backport for [abstractwiki] Enable Abstract Client mode (and on test wiki) (T422657)
- 09:39 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver codfw-k8s # T425377 populating translation memory (dblist: https://phabricator.wikimedia.org/P94307)
- 09:39 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-k8s # T425377 populating translation memory (dblist: https://phabricator.wikimedia.org/P94306)
- 09:35 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pc1017.eqiad.wmnet with reason: host reimage
- 09:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1203: repool after recloning another host
- 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: repool after recloning another host
- 09:29 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on pc1017.eqiad.wmnet with reason: host reimage
- 09:11 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host pc1017.eqiad.wmnet with OS trixie
- 09:10 Msz2001: Deployed private code changes to Suggested Investigations
- 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Reimage to Trixie
- 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Reimage to Trixie
- 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on pc1017.eqiad.wmnet with reason: Reimage to Trixie
- 09:09 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc2017: after reimage to trixie
- 09:08 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2017: after reimage to trixie
- 09:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pc2017.codfw.wmnet with OS trixie
- 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Migration of es1048.eqiad.wmnet completed
- 08:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1048.eqiad.wmnet with OS trixie
- 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: repool after recloning another host
- 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 08:44 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pc2017.codfw.wmnet with reason: host reimage
- 08:40 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1048.eqiad.wmnet with reason: host reimage
- 08:40 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on pc2017.codfw.wmnet with reason: host reimage
- 08:38 dcausse@deploy1003: Finished scap sync-world: Backport for ttmserver-export: pass source language for translation batch IDs (T429479) (duration: 14m 31s)
- 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1048.eqiad.wmnet with reason: host reimage
- 08:32 dcausse@deploy1003: dcausse: Continuing with deployment
- 08:30 dcausse@deploy1003: dcausse: Backport for ttmserver-export: pass source language for translation batch IDs (T429479) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 08:29 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
- 08:28 dcausse@deploy1003: mwscript-k8s job started: purgeList.php --wiki enwiki # purging cache changed images (T414873, T414868)
- 08:28 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
- 08:25 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
- 08:24 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
- 08:24 dcausse@deploy1003: Started scap sync-world: Backport for ttmserver-export: pass source language for translation batch IDs (T429479)
- 08:23 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host pc2017.codfw.wmnet with OS trixie
- 08:21 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
- 08:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: Reimage to Trixie
- 08:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:21 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:21 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2017: Reimage to Trixie
- 08:21 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1048.eqiad.wmnet with OS trixie
- 08:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on pc2017.codfw.wmnet with reason: Reimage to Trixie
- 08:21 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
- 08:17 dcausse@deploy1003: Finished scap sync-world: Backport for shwiki: update wordmark and tagline (T414873), shwiktionary: update logo, wordmark and tagline (T414868) (duration: 56m 56s)
- 08:17 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 08:17 marostegui@cumin1003: dbmaint on pc7@codfw T429178
- 08:17 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1048: Upgrading es1048.eqiad.wmnet
- 08:12 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 08:12 marostegui@cumin1003: dbmaint on pc7@codfw T429178
- 08:12 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: pc7 migration to debian trixie
- 08:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:12 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2017: pc7 migration to debian trixie
- 08:10 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1048: Upgrading es1048.eqiad.wmnet
- 08:10 marostegui@cumin1003: dbmaint on es7@eqiad T429463
- 08:10 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2009.codfw.wmnet with OS bookworm
- 08:05 dcausse@deploy1003: dcausse, vipz: Continuing with deployment
- 08:00 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
- 07:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 07:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2009.codfw.wmnet with reason: host reimage
- 07:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2009.codfw.wmnet with reason: host reimage
- 07:37 dcausse@deploy1003: dcausse, vipz: Backport for shwiki: update wordmark and tagline (T414873), shwiktionary: update logo, wordmark and tagline (T414868) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Moving db1203 to x1 T429562
- 07:32 XioNoX: update pfw policies - T429543
- 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2009.codfw.wmnet with OS bookworm
- 07:20 jmm@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2009.codfw.wmnet with OS bookworm
- 07:20 dcausse@deploy1003: Started scap sync-world: Backport for shwiki: update wordmark and tagline (T414873), shwiktionary: update logo, wordmark and tagline (T414868)
- 07:17 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2009.codfw.wmnet with OS bookworm
- 07:13 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 07:13 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2039: Migration of es2039.codfw.wmnet completed
- 06:51 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Moving db1203 to x1 T429562
- 06:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:34 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:34 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:31 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:31 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:26 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:26 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:24 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1203.eqiad.wmnet
- 06:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1179,1203].eqiad.wmnet with reason: upgrading
- 06:23 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1203: Moving db1203 to x1 T429562
- 06:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1203: Moving db1203 to x1 T429562
- 06:07 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2039: Migration of es2039.codfw.wmnet completed
- 05:54 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2039.codfw.wmnet with OS trixie
- 05:42 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2039.codfw.wmnet with reason: upgrading
- 05:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on es2039.codfw.wmnet with reason: host reimage
- 05:37 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2039.codfw.wmnet with reason: host reimage
- 05:20 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2039.codfw.wmnet with OS trixie
- 05:19 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2039: Upgrading es2039.codfw.wmnet
- 05:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2039: Upgrading es2039.codfw.wmnet
- 05:18 marostegui@cumin1003: dbmaint on es7@codfw T429463
- 05:18 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-21
- 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 12s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-20
- 13:32 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 13:31 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 13:31 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 13:31 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 38s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-19
- 19:21 krinkle@deploy1003: Finished scap sync-world: Backport for Disable ShortUrl on remaining wikis (T107188) (duration: 80m 14s)
- 19:17 krinkle@deploy1003: krinkle: Continuing with deployment
- 18:03 krinkle@deploy1003: krinkle: Backport for Disable ShortUrl on remaining wikis (T107188) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:01 krinkle@deploy1003: Started scap sync-world: Backport for Disable ShortUrl on remaining wikis (T107188)
- 16:22 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2001.codfw.wmnet
- 16:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1023.eqiad.wmnet
- 16:08 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2002.codfw.wmnet
- 16:01 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1023.eqiad.wmnet
- 16:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1022.eqiad.wmnet
- 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1022.eqiad.wmnet
- 15:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1021.eqiad.wmnet
- 15:45 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2002.codfw.wmnet
- 15:44 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1021.eqiad.wmnet
- 15:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1020.eqiad.wmnet
- 15:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1020.eqiad.wmnet
- 15:34 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2001.codfw.wmnet
- 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs2004.codfw.wmnet
- 15:22 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs2004.codfw.wmnet
- 15:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs2003.codfw.wmnet
- 15:17 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs2003.codfw.wmnet
- 15:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs2002.codfw.wmnet
- 15:11 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs2002.codfw.wmnet
- 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs2001.codfw.wmnet
- 14:00 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2009.codfw.wmnet with OS trixie
- 13:44 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2009.codfw.wmnet with reason: host reimage
- 13:41 jmm@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2009.codfw.wmnet with reason: host reimage
- 13:28 jmm@cumin2002: START - Cookbook sre.hosts.reimage for host sretest2009.codfw.wmnet with OS trixie
- 13:02 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs2001.codfw.wmnet
- 13:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs1003.eqiad.wmnet
- 12:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs1003.eqiad.wmnet
- 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs1002.eqiad.wmnet
- 12:51 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs1002.eqiad.wmnet
- 12:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-wdqs1001.eqiad.wmnet
- 12:46 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-wdqs1001.eqiad.wmnet
- 12:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1022.eqiad.wmnet
- 12:32 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1022.eqiad.wmnet
- 12:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2235.codfw.wmnet
- 12:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2235.codfw.wmnet
- 12:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2235.codfw.wmnet
- 12:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2235.codfw.wmnet
- 12:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2234.codfw.wmnet
- 12:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2234.codfw.wmnet
- 12:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2232.codfw.wmnet
- 12:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2232.codfw.wmnet
- 12:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2160.codfw.wmnet
- 12:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2160.codfw.wmnet
- 12:10 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 12:08 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on phab2002.codfw.wmnet with reason: Host Replacement
- 12:05 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:migrateMentorStatusAway.php --wiki=viwiki # T409170
- 12:04 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:MigrateMentorStatusAway --wiki=viwiki # T409170
- 11:33 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 11:23 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:38 moritzm: imported nodejs 24.17.0-1nodesource1 to thirdparty/node24 for trixie-wikimedia
- 10:37 moritzm: imported nodejs 22.23.0-1nodesource1 to thirdparty/node22 for trixie-wikimedia
- 10:33 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2004.codfw.wmnet
- 10:33 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2003.codfw.wmnet
- 10:33 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2002.codfw.wmnet
- 10:33 btullis@puppetserver1001: conftool action : set/pooled=no; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2001.codfw.wmnet
- 10:29 sergi0: Run `MigrateMentorStatusAway` script for all wikis in growthexperiments dblist - T409170
- 10:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1020.eqiad.wmnet
- 10:09 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1020.eqiad.wmnet
- 10:04 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1024
- 10:03 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1024
- 10:03 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023
- 10:03 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023
- 10:03 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1021
- 10:03 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1021
- 10:00 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1024
- 09:59 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1024
- 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1022
- 09:57 btullis@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1022
- 09:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1020
- 09:54 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1020
- 09:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1020.eqiad.wmnet
- 09:36 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1020.eqiad.wmnet
- 07:32 slyngs: Update IDP/SSO to CAS v7.3.7.3
- 07:31 slyngshede@dns1004: END - running authdns-update
- 07:30 slyngshede@dns1004: START - running authdns-update
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:19 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
- 01:18 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
- 01:18 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
- 01:17 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
- 01:17 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
- 01:17 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
- 01:06 ottomata: roll restart eventgate-analytics to pick up stream config change - T427787
2026-06-18
- 23:46 Amir1: ALTER TABLE reading_list_project AUTO_INCREMENT = 882; on wikishared on x1 master (T428002)
- 23:34 rzl@deploy1003: Finished deploy [docker-pkg/deploy@f030aed]: (no justification provided) (duration: 00m 45s)
- 23:33 rzl@deploy1003: Started deploy [docker-pkg/deploy@f030aed]: (no justification provided)
- 23:28 rzl@deploy1003: Finished deploy [docker-pkg/deploy@f030aed]: (no justification provided) (duration: 00m 26s)
- 23:27 rzl@deploy1003: Started deploy [docker-pkg/deploy@f030aed]: (no justification provided)
- 23:03 rzl: rzl@apt1002:~$ sudo -i reprepro -C main include trixie-wikimedia /home/rzl/httpbb/trixie/httpbb_0.0.5-1+deb13u1_amd64.changes # T427899
- 22:52 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Re-enable for mcrundo (T427612) (duration: 07m 25s)
- 22:47 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 22:46 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Re-enable for mcrundo (T427612) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:44 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Re-enable for mcrundo (T427612)
- 21:29 maryum: Deployed security fix for T428833
- 21:14 jdlrobson@deploy1003: Finished scap sync-world: Backport for Prevent surveys being automatically added to non-Wikipedias (T393436) (duration: 07m 54s)
- 21:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
- 21:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
- 21:09 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 21:08 jdlrobson@deploy1003: jdlrobson: Backport for Prevent surveys being automatically added to non-Wikipedias (T393436) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:06 jdlrobson@deploy1003: Started scap sync-world: Backport for Prevent surveys being automatically added to non-Wikipedias (T393436)
- 20:12 dani@deploy1003: Finished scap sync-world: Backport for Deploy English Wikipedia Mobile App Survey (T428876) (duration: 08m 20s)
- 20:08 dani@deploy1003: dani: Continuing with deployment
- 20:06 dani@deploy1003: dani: Backport for Deploy English Wikipedia Mobile App Survey (T428876) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:04 dani@deploy1003: Started scap sync-world: Backport for Deploy English Wikipedia Mobile App Survey (T428876)
- 19:11 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*
- 19:09 cdobbins@dns1004: END - running authdns-update
- 19:08 cdobbins@dns1004: START - running authdns-update
- 19:07 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update
- 19:05 cdobbins@dns1004: END - running authdns-update
- 19:04 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2002.codfw.wmnet with reason: Host Replacement
- 19:03 cdobbins@dns1004: START - running authdns-update
- 19:01 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns7002.wikimedia.org
- 19:01 cdobbins@cumin2002: START - Cookbook sre.hosts.remove-downtime for dns7002.wikimedia.org
- 18:54 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm
- 18:39 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.7 refs T423916
- 18:37 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 18:34 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 18:33 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 18:31 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 18:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 18:28 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 18:27 swfrench-wmf: (eqiad) kubectl delete pod coredns-54cdd9bdf-6hwb5 -n kube-system - T429156
- 18:27 swfrench-wmf: (eqiad) kubectl delete pod coredns-54cdd9bdf-6n4ps -n kube-system - T429156
- 18:26 jhuneidi@deploy1003: Finished scap sync-world: Backport for SpecialSpecialPages: Guard against special pages with no content-language alias (T429584) (duration: 08m 46s)
- 18:21 jhuneidi@deploy1003: jhuneidi, jforrester: Continuing with deployment
- 18:19 jhuneidi@deploy1003: jhuneidi, jforrester: Backport for SpecialSpecialPages: Guard against special pages with no content-language alias (T429584) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:17 jhuneidi@deploy1003: Started scap sync-world: Backport for SpecialSpecialPages: Guard against special pages with no content-language alias (T429584)
- 18:09 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
- 18:04 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
- 17:37 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm
- 16:28 zabe@deploy1003: Finished scap sync-world: Backport for Add script to fix fr_archive_name drifts (T428406) (duration: 06m 46s)
- 16:24 zabe@deploy1003: zabe: Continuing with deployment
- 16:24 zabe@deploy1003: zabe: Backport for Add script to fix fr_archive_name drifts (T428406) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 16:22 zabe@deploy1003: Started scap sync-world: Backport for Add script to fix fr_archive_name drifts (T428406)
- 15:55 zabe@deploy1003: Finished scap sync-world: Backport for LocalFileMoveBatch: Also update fr_archive_name when moving file (T428406) (duration: 06m 49s)
- 15:51 zabe@deploy1003: zabe: Continuing with deployment
- 15:51 zabe@deploy1003: zabe: Backport for LocalFileMoveBatch: Also update fr_archive_name when moving file (T428406) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:49 zabe@deploy1003: Started scap sync-world: Backport for LocalFileMoveBatch: Also update fr_archive_name when moving file (T428406)
- 15:21 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
- 15:21 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
- 15:21 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 15:21 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 15:08 elukey@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox
- 15:08 elukey@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox
- 15:04 cscott@deploy1003: Finished scap sync-world: Backport for Check that data-parsoid is an array before accessing it as such (T429582) (duration: 11m 17s)
- 15:00 cscott@deploy1003: ihurbain, cscott: Continuing with deployment
- 14:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2002.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:57 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2002.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:55 cscott@deploy1003: ihurbain, cscott: Backport for Check that data-parsoid is an array before accessing it as such (T429582) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:53 cscott@deploy1003: Started scap sync-world: Backport for Check that data-parsoid is an array before accessing it as such (T429582)
- 14:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:51 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:51 ayounsi@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) homer to cumin2003.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:46 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: trixie homer deploy - ayounsi@cumin1003
- 14:42 moritzm: installing zsh updates from Bookworm point release
- 14:37 brouberol@dns1004: END - running authdns-update
- 14:35 brouberol@dns1004: START - running authdns-update
- 14:27 jgreen@dns1004: END - running authdns-update
- 14:25 jgreen@dns1004: START - running authdns-update
- 14:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dbproxy2007.codfw.wmnet
- 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for dbproxy2007.codfw.wmnet
- 14:21 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dbproxy2008.codfw.wmnet
- 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for dbproxy2008.codfw.wmnet
- 14:20 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2160.codfw.wmnet
- 14:20 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2160.codfw.wmnet
- 14:19 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2235.codfw.wmnet
- 14:19 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2235.codfw.wmnet
- 14:14 Msz2001: Finished deploying private code change
- 14:08 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2235.codfw.wmnet with reason: Reboots T426633
- 14:08 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbproxy2008.codfw.wmnet with reason: Reboots T426633
- 14:08 moritzm: installing unbound security updates
- 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2234.codfw.wmnet
- 14:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2234.codfw.wmnet
- 14:00 tgr_: UTC afternoon deploys done
- 14:00 tgr@deploy1003: Finished scap sync-world: Backport for Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495), Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495) (duration: 11m 51s)
- 13:56 tgr@deploy1003: tgr: Continuing with deployment
- 13:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2234.codfw.wmnet with reason: Reboots T426633
- 13:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2160.codfw.wmnet with reason: Reboots T426633
- 13:54 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbproxy2007.codfw.wmnet with reason: Reboots T426633
- 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dbproxy2005.codfw.wmnet
- 13:52 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for dbproxy2005.codfw.wmnet
- 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2232.codfw.wmnet
- 13:51 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2232.codfw.wmnet
- 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2160.codfw.wmnet
- 13:51 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2160.codfw.wmnet
- 13:50 tgr@deploy1003: tgr: Backport for Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495), Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:48 tgr@deploy1003: Started scap sync-world: Backport for Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495), Fix CentralAuthPostLoginRedirect type parameter on token loss (T429495)
- 13:46 tgr@deploy1003: Finished scap sync-world: Backport for magwiki: add wordmark, metanamespace, sitename and timezone (T428279), stream: webrequest.page_trending.dev0 (T429588) (duration: 08m 15s)
- 13:42 tgr@deploy1003: javiermonton, tgr, anzx: Continuing with deployment
- 13:41 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.changedisk (exit_code=0) for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 13:40 tgr@deploy1003: javiermonton, tgr, anzx: Backport for magwiki: add wordmark, metanamespace, sitename and timezone (T428279), stream: webrequest.page_trending.dev0 (T429588) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:38 tgr@deploy1003: Started scap sync-world: Backport for magwiki: add wordmark, metanamespace, sitename and timezone (T428279), stream: webrequest.page_trending.dev0 (T429588)
- 13:38 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2160.codfw.wmnet with reason: Reboots T426633
- 13:38 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2232.codfw.wmnet with reason: Reboots T426633
- 13:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbproxy2005.codfw.wmnet with reason: Reboots T426633
- 13:33 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 13:30 ladsgroup@deploy1003: Finished scap sync-world: Backport for REST: Adjust key of Reading Lists OpenAPI spec in RestSandboxSpecs (T422771) (duration: 06m 56s)
- 13:26 ladsgroup@deploy1003: ladsgroup, bpirkle: Continuing with deployment
- 13:25 ladsgroup@deploy1003: ladsgroup, bpirkle: Backport for REST: Adjust key of Reading Lists OpenAPI spec in RestSandboxSpecs (T422771) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:23 ladsgroup@deploy1003: Started scap sync-world: Backport for REST: Adjust key of Reading Lists OpenAPI spec in RestSandboxSpecs (T422771)
- 13:23 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.changedisk (exit_code=0) for changing disk type of testvm2005.codfw.wmnet to drbd
- 13:21 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of testvm2005.codfw.wmnet to drbd
- 13:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for EventStreamConfig: add stream for WDQS V2 external/internal queries. (T429380) (duration: 10m 55s)
- 13:14 ladsgroup@deploy1003: ladsgroup, lerickson: Continuing with deployment
- 13:10 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.changedisk (exit_code=99) for changing disk type of testvm2005.codfw.wmnet to drbd
- 13:10 ladsgroup@deploy1003: ladsgroup, lerickson: Backport for EventStreamConfig: add stream for WDQS V2 external/internal queries. (T429380) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:08 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of testvm2005.codfw.wmnet to drbd
- 13:08 fabfur: deploying new haproxykafka on A:cp to parse for x_provenance (T427068)
- 13:08 ladsgroup@deploy1003: Started scap sync-world: Backport for EventStreamConfig: add stream for WDQS V2 external/internal queries. (T429380)
- 13:07 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.changedisk (exit_code=0) for changing disk type of testvm2005.codfw.wmnet to plain
- 13:05 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of testvm2005.codfw.wmnet to plain
- 13:03 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2001.codfw.wmnet
- 13:03 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2002.codfw.wmnet
- 13:03 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2003.codfw.wmnet
- 13:03 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2004.codfw.wmnet
- 13:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis magwiki in section s5
- 13:00 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2004.codfw.wmnet
- 13:00 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2003.codfw.wmnet
- 13:00 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2002.codfw.wmnet
- 13:00 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-wdqs2001.codfw.wmnet
- 12:56 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.changedisk (exit_code=99) for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 12:39 fabfur: upgrade haproxykafka on cp1111 to test for new x-provenance field (T427068)
- 12:36 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 12:35 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
- 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis magwiki in section s5
- 12:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
- 12:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis magwiki in section s5
- 12:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for TranslatePage: Cast to string before using htmlspecialchars (T429459), TranslatePage: Cast to string before using htmlspecialchars (T429459) (duration: 17m 49s)
- 12:29 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis magwiki in section s5
- 12:27 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 12:16 dreamyjazz@deploy1003: dreamyjazz: Backport for TranslatePage: Cast to string before using htmlspecialchars (T429459), TranslatePage: Cast to string before using htmlspecialchars (T429459) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:14 dreamyjazz@deploy1003: Started scap sync-world: Backport for TranslatePage: Cast to string before using htmlspecialchars (T429459), TranslatePage: Cast to string before using htmlspecialchars (T429459)
- 11:10 atsukoito: atsuko updated charlie to 0.0.19 https://w.wiki/RPKN
- 10:37 jmm@cumin2002: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99)
- 10:37 jmm@cumin2002: START - Cookbook sre.puppet.disable-merges
- 10:24 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Recompute blocked-edit risk score block IDs server-side (T428394) (duration: 12m 13s)
- 10:19 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 10:14 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Recompute blocked-edit risk score block IDs server-side (T428394) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:11 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Recompute blocked-edit risk score block IDs server-side (T428394)
- 10:05 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
- 10:05 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
- 10:01 fabfur@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Change provenance var context - fabfur@cumin1003 - T427068"
- 10:01 fabfur@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Change provenance var context - fabfur@cumin1003 - T427068
- 10:00 fabfur@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Change provenance var context - fabfur@cumin1003 - T427068
- 10:00 fabfur@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Change provenance var context - fabfur@cumin1003 - T427068"
- 09:59 kharlan@deploy1003: Finished scap sync-world: Backport for CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481), CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481) (duration: 08m 10s)
- 09:55 kharlan@deploy1003: kharlan: Continuing with deployment
- 09:54 kharlan@deploy1003: kharlan: Backport for CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481), CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:51 kharlan@deploy1003: Started scap sync-world: Backport for CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481), CaptchaScoreHooks: Log risk score for every non-exempt edit (T429481)
- 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
- 09:33 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
- 09:33 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
- 09:32 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
- 09:11 moritzm: installing apache2 security updates
- 08:55 jelto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
- 08:53 jelto@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
- 08:53 jelto@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
- 08:51 jelto@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
- 08:51 jelto@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
- 08:51 jelto@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
- 08:35 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 08:34 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 08:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 08:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 08:22 jelto@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
- 08:21 jelto@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
- 08:20 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
- 08:19 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
- 08:05 moritzm: regenerate pbuilder environments on build2001 to use deb.debian.org T416707
- 08:02 moritzm: uploaded wmf-laptop 1.0.6 to component/wmf-laptop on apt.wikimedia.org
- 08:01 moritzm: regenerate pbuilder environments on build2002 to use deb.debian.org T416707
- 06:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 06:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Migration of es2040.codfw.wmnet completed
- 06:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Migration of es2040.codfw.wmnet completed
- 05:52 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2040.codfw.wmnet with OS trixie
- 05:41 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
- 05:41 marostegui@cumin1003: Removing db1224 from zarcillo T429561
- 05:41 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1224.eqiad.wmnet
- 05:41 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 05:41 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1224.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 05:40 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1224.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
- 05:36 marostegui@cumin1003: START - Cookbook sre.dns.netbox
- 05:35 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2040.codfw.wmnet with reason: host reimage
- 05:31 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2040.codfw.wmnet with reason: host reimage
- 05:31 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1224.eqiad.wmnet
- 05:30 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
- 05:27 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1224 from dbctl T429561', diff saved to https://phabricator.wikimedia.org/P94269 and previous config saved to /var/cache/conftool/dbconfig/20260618-052737-marostegui.json
- 05:14 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2040.codfw.wmnet with OS trixie
- 05:13 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2040: Upgrading es2040.codfw.wmnet
- 05:13 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2040: Upgrading es2040.codfw.wmnet
- 05:12 marostegui@cumin1003: dbmaint on es7@codfw T429463
- 05:12 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for Update interwiki map (T428266) (duration: 06m 55s)
- 01:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 01:14 ladsgroup@deploy1003: ladsgroup: Backport for Update interwiki map (T428266) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 01:12 ladsgroup@deploy1003: Started scap sync-world: Backport for Update interwiki map (T428266)
- 00:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for Activate magwiki (T428266) (duration: 07m 25s)
- 00:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 00:42 ladsgroup@deploy1003: ladsgroup: Backport for Activate magwiki (T428266) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 00:40 ladsgroup@deploy1003: Started scap sync-world: Backport for Activate magwiki (T428266)
- 00:33 ladsgroup@deploy1003: Finished scap sync-world: Backport for Init magwiki (T428266) (duration: 07m 14s)
- 00:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
- 00:28 ladsgroup@deploy1003: ladsgroup: Backport for Init magwiki (T428266) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 00:26 ladsgroup@deploy1003: Started scap sync-world: Backport for Init magwiki (T428266)
2026-06-17
- 23:26 egardner@deploy1003: Finished scap sync-world: Backport for Enable beta mobile MMV on Wikipedias (T426775) (duration: 06m 46s)
- 23:22 egardner@deploy1003: egardner: Continuing with deployment
- 23:21 egardner@deploy1003: egardner: Backport for Enable beta mobile MMV on Wikipedias (T426775) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 23:19 egardner@deploy1003: Started scap sync-world: Backport for Enable beta mobile MMV on Wikipedias (T426775)
- 23:17 egardner@deploy1003: Finished scap sync-world: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193) (duration: 06m 55s)
- 23:14 mutante: gerrit2002 - unlink /srv/gerrit/site_path/review_site/logs/logs (T425667)
- 23:12 egardner@deploy1003: egardner: Continuing with deployment
- 23:12 egardner@deploy1003: egardner: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 23:10 egardner@deploy1003: Started scap sync-world: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193)
- 23:04 egardner@deploy1003: Finished scap sync-world: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193) (duration: 12m 31s)
- 22:57 egardner@deploy1003: egardner: Continuing with deployment
- 22:56 egardner@deploy1003: egardner: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:52 egardner@deploy1003: Started scap sync-world: Backport for Image Browsing: fix transparent images in carousel (T429047), MMV Beta Viewer: Make in-flight image downloads abortable (T429193), MMV Beta Viewer: Delay the loading indicator on quick navigation (T429193)
- 22:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for Donor Delight Badge: Add accessible label and hide popover from AT (T427313) (duration: 31m 01s)
- 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 22:31 jdlrobson@deploy1003: jdlrobson: Backport for Donor Delight Badge: Add accessible label and hide popover from AT (T427313) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 22:14 jdlrobson@deploy1003: Started scap sync-world: Backport for Donor Delight Badge: Add accessible label and hide popover from AT (T427313)
- 21:52 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:52 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:29 ecarg@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:29 ecarg@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:29 ecarg@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:28 ecarg@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:27 ecarg@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 21:27 ecarg@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 21:23 ecarg@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:22 ecarg@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:22 ecarg@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:21 ecarg@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:20 ecarg@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 21:20 ecarg@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 21:15 ecarg@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 21:12 ecarg@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 21:12 ecarg@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 21:09 ecarg@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 21:06 ecarg@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 21:05 ecarg@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 21:02 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
- 21:02 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
- 20:45 cdobbins@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dns7002.wikimedia.org with reason: bird.service keeps failing
- 20:41 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp
- 20:41 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie
- 20:36 sbisson@deploy1003: Finished scap sync-world: Backport for Enable ULS v2 on group1 wikis (duration: 08m 26s)
- 20:31 sbisson@deploy1003: sbisson, abi: Continuing with deployment
- 20:29 sbisson@deploy1003: sbisson, abi: Backport for Enable ULS v2 on group1 wikis synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:27 sbisson@deploy1003: Started scap sync-world: Backport for Enable ULS v2 on group1 wikis
- 20:17 sgimeno@deploy1003: Finished scap sync-world: Backport for migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170), migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170) (duration: 06m 55s)
- 20:13 sgimeno@deploy1003: sgimeno: Continuing with deployment
- 20:12 sgimeno@deploy1003: sgimeno: Backport for migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170), migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:11 sgimeno@deploy1003: Started scap sync-world: Backport for migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170), migrateMentorStatusAway: Return SIMULATED for all dry-run executions (T409170)
- 19:44 jgreen@dns1005: END - running authdns-update
- 19:42 jgreen@dns1005: START - running authdns-update
- 19:31 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs5005*} and A:liberica (T428229)
- 19:30 brett@cumin2002: START - Cookbook sre.loadbalancer.admin config_reloading P{lvs5005*} and A:liberica (T428229)
- 19:16 jhuneidi@deploy1003: Finished scap sync-world: wmf.7 to group 1 (Take 2) (duration: 07m 01s)
- 19:16 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp and not P{cp7001.magru.wmnet} and A:cp
- 19:10 jhuneidi@deploy1003: Started scap sync-world: wmf.7 to group 1 (Take 2)
- 19:08 jhuneidi@deploy1003: Finished scap sync-world: Attempt to roll wmf.7 to group 1 (duration: 07m 24s)
- 19:01 jhuneidi@deploy1003: Started scap sync-world: Attempt to roll wmf.7 to group 1
- 19:00 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudcontrol1008-dev.eqiad.wmnet
- 19:00 andrew@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:00 andrew@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudcontrol1008-dev.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - andrew@cumin2002"
- 18:59 andrew@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudcontrol1008-dev.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - andrew@cumin2002"
- 18:52 andrew@cumin2002: START - Cookbook sre.dns.netbox
- 18:46 andrew@cumin2002: START - Cookbook sre.hosts.decommission for hosts cloudcontrol1008-dev.eqiad.wmnet
- 18:24 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6011.*
- 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp6011.drmrs.wmnet
- 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for cp6011.drmrs.wmnet
- 18:19 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6011.drmrs.wmnet with reason: ats restart, continuing from failed cookbook run
- 18:17 brett: commit new lvs5005 IP address to cr2-eqsin.wikimedia.org,cr3-eqsin.wikimedia.org
- 18:16 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cp6011.drmrs.wmnet
- 18:07 brett@cumin2002: START - Cookbook sre.hosts.reboot-single for host cp6011.drmrs.wmnet
- 18:07 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6011.*
- 17:41 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host lvs5005.eqsin.wmnet with OS bookworm
- 17:20 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs5005.eqsin.wmnet with reason: host reimage
- 17:16 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on lvs5005.eqsin.wmnet with reason: host reimage
- 17:06 mutante: contint1003 - even with gerrit:1301416 jenkins was STILL restarted :/ - stopping it manually and puppet - debugging - T418521
- 17:03 mutante: contint1003 - re-enabling puppet - checking it does NOT start jenkins - also see gerrit:1297236 and gerrit:1301416 - T418521
- 16:51 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 16:51 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 16:49 brett@cumin2002: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp
- 16:48 brett@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host lvs5005
- 16:48 brett@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host lvs5005
- 16:48 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 16:47 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
- 16:47 brett@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host lvs5005
- 16:47 brett@cumin2002: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) lvs5005.eqsin.wmnet 6.0.132.10.in-addr.arpa 6.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 16:47 brett@cumin2002: START - Cookbook sre.dns.wipe-cache lvs5005.eqsin.wmnet 6.0.132.10.in-addr.arpa 6.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 16:45 brett@cumin2002: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) lvs5005.eqsin.wmnet 6.0.132.10.in-addr.arpa 6.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 16:45 brett@cumin2002: START - Cookbook sre.dns.wipe-cache lvs5005.eqsin.wmnet 6.0.132.10.in-addr.arpa 6.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 16:45 brett@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:45 brett@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host lvs5005 - brett@cumin2002"
- 16:45 brett@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host lvs5005 - brett@cumin2002"
- 16:45 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 16:45 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
- 16:39 brett@cumin2002: START - Cookbook sre.dns.netbox
- 16:16 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1078.eqiad.wmnet with OS trixie
- 16:16 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 16:16 brett@cumin2002: START - Cookbook sre.hosts.move-vlan for host lvs5005
- 16:16 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 16:15 brett@cumin2002: START - Cookbook sre.hosts.reimage for host lvs5005.eqsin.wmnet with OS bookworm
- 16:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging1007.eqiad.wmnet with OS trixie
- 16:15 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 16:11 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 16:02 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs5005.eqsin.wmnet} and A:liberica
- 16:02 brett@cumin2002: START - Cookbook sre.loadbalancer.admin depooling P{lvs5005.eqsin.wmnet} and A:liberica
- 16:00 brett@cumin2002: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp and not P{cp7001.magru.wmnet} and A:cp
- 15:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1078.eqiad.wmnet with reason: host reimage
- 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging1007.eqiad.wmnet with reason: host reimage
- 15:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Migration of es2048.codfw.wmnet completed
- 15:53 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1078.eqiad.wmnet with reason: host reimage
- 15:47 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging1007.eqiad.wmnet with reason: host reimage
- 15:46 moritzm: installing python-ldap security updates
- 15:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1078.eqiad.wmnet with OS trixie
- 15:30 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 15:27 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 15:26 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1007.eqiad.wmnet with OS trixie
- 15:08 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Migration of es2048.codfw.wmnet completed
- 15:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 15:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
- 15:03 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp1004.eqiad.wmnet with OS trixie
- 15:02 aokoth@deploy1003: Finished deploy [phabricator/deployment@a640ed9]: deploy phab (duration: 01m 24s)
- 15:00 aokoth@deploy1003: Started deploy [phabricator/deployment@a640ed9]: deploy phab
- 14:59 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
- 14:57 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2048.codfw.wmnet with OS trixie
- 14:56 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
- 14:44 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp1004.eqiad.wmnet with reason: host reimage
- 14:40 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2048.codfw.wmnet with reason: host reimage
- 14:35 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp1004.eqiad.wmnet with reason: host reimage
- 14:33 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2048.codfw.wmnet with reason: host reimage
- 14:28 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
- 14:26 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for Add Wikidata configuration for WikiProject links (T422935 T422936) (duration: 07m 49s)
- 14:22 lucaswerkmeister-wmde@deploy1003: audreypenven, lucaswerkmeister-wmde: Continuing with deployment
- 14:21 cjd91: depooling dns7002 to attempt reimage to trixie
- 14:20 lucaswerkmeister-wmde@deploy1003: audreypenven, lucaswerkmeister-wmde: Backport for Add Wikidata configuration for WikiProject links (T422935 T422936) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:19 blake@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp1004.eqiad.wmnet with OS trixie
- 14:19 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.*
- 14:18 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for Add Wikidata configuration for WikiProject links (T422935 T422936)
- 14:17 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2048.codfw.wmnet with OS trixie
- 14:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
- 14:17 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
- 14:17 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
- 14:16 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
- 14:16 ecarg@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 14:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2048: Upgrading es2048.codfw.wmnet
- 14:13 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2048: Upgrading es2048.codfw.wmnet
- 14:13 elukey: add basic Kafka ACLs for anonymous to logging-eqiad - T425528
- 14:13 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 14:13 Lucas_WMDE: UTC afternoon backport+config window done
- {{safesubst:SAL entry|1=14:13 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for ULS rewrite: Lock body scroll when open on mobile, ULS rewrite: Fix settings dialog width and field sizing (T416512), ULS rewrite: Show variants even when no languages are available (T426532), ULS rewrite: Capture trigger element before async module load (T429145), [[gerr}}
- 14:12 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs-test1001.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1003.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1002.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1001.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs-test1001.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1003.eqiad.wmnet
- 14:12 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1002.eqiad.wmnet
- 14:11 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs1001.eqiad.wmnet
- 14:11 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-wdqs*.eqiad.wmnet
- 14:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, abi: Continuing with deployment
- 14:06 ecarg@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 14:01 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
- 14:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
- 13:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs2003.codfw.wmnet with OS bookworm
- 13:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 13:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs2004.codfw.wmnet with OS bookworm
- 13:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- {{safesubst:SAL entry|1=13:55 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, abi: Backport for ULS rewrite: Lock body scroll when open on mobile, ULS rewrite: Fix settings dialog width and field sizing (T416512), ULS rewrite: Show variants even when no languages are available (T426532), ULS rewrite: Capture trigger element before async module load (T429145), [[ge}}
- {{safesubst:SAL entry|1=13:53 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for ULS rewrite: Lock body scroll when open on mobile, ULS rewrite: Fix settings dialog width and field sizing (T416512), ULS rewrite: Show variants even when no languages are available (T426532), ULS rewrite: Capture trigger element before async module load (T429145), [[gerri}}
- 13:52 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
- 13:51 jmm@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
- 13:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest1005.eqiad.wmnet
- 13:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest1005.eqiad.wmnet
- 13:47 papaul: mgmt interface change on mr-codfw
- 13:46 pt1979@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mr1-codfw with reason: mgmt interface change
- 13:45 pt1979@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mr1-codfw with reason: switch refresh
- 13:42 jmm@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
- 13:42 jmm@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
- 13:33 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for Add Wikidata configuration for WikiProject links (T422935), Add instance-of WikiProject links for paintings and elections (T422936) (duration: 08m 14s)
- 13:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp1006.eqiad.wmnet with OS trixie
- 13:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudcephosd1016
- 13:31 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudcephosd1016
- 13:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudvirt1061
- 13:31 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudvirt1061
- 13:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudvirt1069
- 13:31 lucaswerkmeister-wmde@deploy1003: sadiyamohammed13, lucaswerkmeister-wmde: Rolling back deployment
- 13:31 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudvirt1069
- 13:30 cmooney@cumin1003: END (PASS) - Cookbook sre.network.cloud-host (exit_code=0) for host cloudvirt1068
- 13:30 cmooney@cumin1003: START - Cookbook sre.network.cloud-host for host cloudvirt1068
- 13:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp1005.eqiad.wmnet with OS trixie
- 13:27 lucaswerkmeister-wmde@deploy1003: sadiyamohammed13, lucaswerkmeister-wmde: Backport for Add Wikidata configuration for WikiProject links (T422935), Add instance-of WikiProject links for paintings and elections (T422936) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:25 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for Add Wikidata configuration for WikiProject links (T422935), Add instance-of WikiProject links for paintings and elections (T422936)
- 13:24 jmm@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
- 13:23 jmm@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
- 13:14 dani@deploy1003: Finished scap sync-world: Backport for Add English Wikipedia Mobile App Survey (T428876) (duration: 07m 53s)
- 13:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp1006.eqiad.wmnet with reason: host reimage
- 13:11 klausman@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-reboot (exit_code=0) rolling reboot on A:ml-cache-codfw
- 13:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp1005.eqiad.wmnet with reason: host reimage
- 13:10 klausman@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-reboot (exit_code=0) rolling reboot on A:ml-cache-eqiad
- 13:10 dani@deploy1003: dani: Continuing with deployment
- 13:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1045: repool after upgrade
- 13:08 dani@deploy1003: dani: Backport for Add English Wikipedia Mobile App Survey (T428876) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:07 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp1006.eqiad.wmnet with reason: host reimage
- 13:06 dani@deploy1003: Started scap sync-world: Backport for Add English Wikipedia Mobile App Survey (T428876)
- 13:06 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp1005.eqiad.wmnet with reason: host reimage
- 13:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:52 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-gp1006
- 12:52 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp1006
- 12:51 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp1006
- 12:51 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp1006.eqiad.wmnet 182.48.64.10.in-addr.arpa 2.8.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 12:51 blake@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp1006.eqiad.wmnet 182.48.64.10.in-addr.arpa 2.8.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 12:51 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:51 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-gp1005
- 12:51 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp1005
- 12:49 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp1005
- 12:49 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp1005.eqiad.wmnet 126.32.64.10.in-addr.arpa 6.2.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 12:49 blake@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp1005.eqiad.wmnet 126.32.64.10.in-addr.arpa 6.2.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 12:49 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:49 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-gp1005 - blake@cumin1003"
- 12:49 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-gp1005 - blake@cumin1003"
- 12:48 blake@cumin1003: START - Cookbook sre.dns.netbox
- 12:45 klausman@cumin1003: START - Cookbook sre.cassandra.roll-reboot rolling reboot on A:ml-cache-codfw
- 12:45 klausman@cumin1003: START - Cookbook sre.cassandra.roll-reboot rolling reboot on A:ml-cache-eqiad
- 12:43 blake@cumin1003: START - Cookbook sre.dns.netbox
- 12:41 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc-gp1006
- 12:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 12:41 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:ml-cache-codfw: Security updates (T426585) - klausman@cumin1003
- 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:ml-cache-eqiad: Security updates (T426585) - klausman@cumin1003
- 12:41 blake@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp1006.eqiad.wmnet with OS trixie
- 12:41 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc-gp1005
- 12:40 blake@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp1005.eqiad.wmnet with OS trixie
- 12:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2004.codfw.wmnet with reason: host reimage
- 12:37 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Migration of db1163.eqiad.wmnet completed
- 12:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2003.codfw.wmnet with reason: host reimage
- 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:32 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
- 12:32 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
- 12:32 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
- 12:32 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
- 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs2004.codfw.wmnet with reason: host reimage
- 12:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs2003.codfw.wmnet with reason: host reimage
- 12:24 klausman@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:ml-cache-codfw: Security updates (T426585) - klausman@cumin1003
- 12:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1045: repool after upgrade
- 12:23 klausman@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:ml-cache-eqiad: Security updates (T426585) - klausman@cumin1003
- 12:22 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1045.eqiad.wmnet with OS trixie
- 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: host reimage
- 12:19 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 12:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2004.codfw.wmnet with OS bookworm
- 12:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2003.codfw.wmnet with OS bookworm
- 12:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: host reimage
- 12:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2002.codfw.wmnet with reason: host reimage
- 12:08 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:07 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 12:07 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 12:07 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 12:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:07 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 12:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:05 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs2002.codfw.wmnet with reason: host reimage
- 12:04 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1045.eqiad.wmnet with reason: host reimage
- 12:03 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:03 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:03 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2044: repool after maintenance es2044
- 12:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 12:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 12:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:00 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:00 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1045.eqiad.wmnet with reason: host reimage
- 11:55 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 11:55 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 11:55 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 11:54 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 11:51 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 11:51 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Migration of db1163.eqiad.wmnet completed
- 11:44 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1045.eqiad.wmnet with OS trixie
- 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1045: Upgrading es1045.eqiad.wmnet
- 11:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1045: Upgrading es1045.eqiad.wmnet
- 11:42 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 11:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1163.eqiad.wmnet with OS trixie
- 11:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2002.codfw.wmnet with reason: host reimage
- 11:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs2002.codfw.wmnet with reason: host reimage
- 11:27 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 11:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1191.eqiad.wmnet with reason: upgrading
- 11:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 11:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1163.eqiad.wmnet with reason: host reimage
- 11:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1172.eqiad.wmnet with reason: upgrading
- 11:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.dhcp (exit_code=0) for host dse-k8s-wdqs2001.codfw.wmnet
- 11:21 marostegui@cumin1003: DONE (ERROR) - Cookbook sre.hosts.downtime (exit_code=97) for 1:00:00 on db1171.eqiad.wmnet with reason: upgrading
- 11:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1190.eqiad.wmnet with reason: upgrading
- 11:18 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1163.eqiad.wmnet with reason: host reimage
- 11:18 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 11:17 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2044: repool after maintenance es2044
- 11:17 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 11:16 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2044.codfw.wmnet with OS trixie
- 11:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs1003.eqiad.wmnet with OS bookworm
- 11:12 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 11:11 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
- 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
- 11:10 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 11:09 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
- 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
- 11:08 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Migration of es1038.eqiad.wmnet completed
- 11:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1163.eqiad.wmnet with OS trixie
- 11:02 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
- 11:02 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
- 11:01 moritzm: The Debian mirror on mirrors.wikimedia.org has been disabled T416707
- 11:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1163: Upgrading db1163.eqiad.wmnet
- 10:59 btullis@cumin1003: START - Cookbook sre.hosts.dhcp for host dse-k8s-wdqs2001.codfw.wmnet
- 10:59 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1163: Upgrading db1163.eqiad.wmnet
- 10:59 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:59 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2044.codfw.wmnet with reason: host reimage
- 10:53 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2044.codfw.wmnet with reason: host reimage
- 10:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs1003.eqiad.wmnet with reason: host reimage
- 10:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Migration of db2203.codfw.wmnet completed
- 10:43 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs1003.eqiad.wmnet with reason: host reimage
- 10:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 10:37 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2044.codfw.wmnet with OS trixie
- 10:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2044: Upgrading es2044.codfw.wmnet
- 10:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2044: Upgrading es2044.codfw.wmnet
- 10:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
- 10:35 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
- 10:35 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply
- 10:34 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply
- 10:34 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/ratelimit: apply
- 10:34 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/ratelimit: apply
- 10:31 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs1003.eqiad.wmnet with OS bookworm
- 10:29 moritzm: installing git-lfs security updates
- 10:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2002.codfw.wmnet with OS bookworm
- 10:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs1002.eqiad.wmnet with OS bookworm
- 10:28 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 10:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Migration of es1038.eqiad.wmnet completed
- 10:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 10:21 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 10:17 claime: cumin -x 'A:swift-fe' "enable-puppet 'Disabling puppet for ratelimit deploy - cgoubert'"
- 10:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1038.eqiad.wmnet with OS trixie
- 10:12 claime: cumin -x 'A:swift-fe' "disable-puppet 'Disabling puppet for ratelimit deploy - cgoubert'"
- 10:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 10:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich-next: apply
- 10:09 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich-next: apply
- 10:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs1002.eqiad.wmnet with reason: host reimage
- 10:02 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Migration of db2203.codfw.wmnet completed
- 10:00 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs1002.eqiad.wmnet with reason: host reimage
- 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1038.eqiad.wmnet with reason: host reimage
- 09:54 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1038.eqiad.wmnet with reason: host reimage
- 09:52 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2203.codfw.wmnet with OS trixie
- 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2045: repool after maintenance es2045
- 09:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs1002.eqiad.wmnet with OS bookworm
- 09:47 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Remove config for VE and DT enable (T428883), Drop $wgDiscussionToolsHCaptchaRequiredForAllEdits (T428883) (duration: 15m 32s)
- 09:41 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 09:39 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs1002.eqiad.wmnet with OS bookworm
- 09:38 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1038.eqiad.wmnet with OS trixie
- 09:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: Upgrading es1038.eqiad.wmnet
- 09:38 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Remove config for VE and DT enable (T428883), Drop $wgDiscussionToolsHCaptchaRequiredForAllEdits (T428883) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:37 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: Upgrading es1038.eqiad.wmnet
- 09:37 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 09:37 marostegui@dns1004: END - running authdns-update
- 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Set es6 eqiad back to read-write - T429436', diff saved to https://phabricator.wikimedia.org/P94226 and previous config saved to /var/cache/conftool/dbconfig/20260617-093559-marostegui.json
- 09:35 marostegui@dns1004: START - running authdns-update
- 09:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es1038 T429436', diff saved to https://phabricator.wikimedia.org/P94225 and previous config saved to /var/cache/conftool/dbconfig/20260617-093513-marostegui.json
- 09:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2203.codfw.wmnet with reason: host reimage
- 09:33 marostegui@cumin1003: dbctl commit (dc=all): 'Promote es1037 to es6 primary T429436', diff saved to https://phabricator.wikimedia.org/P94224 and previous config saved to /var/cache/conftool/dbconfig/20260617-093310-marostegui.json
- 09:32 marostegui: Starting es6 eqiad failover from es1038 to es1037 - T429436
- 09:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Remove config for VE and DT enable (T428883), Drop $wgDiscussionToolsHCaptchaRequiredForAllEdits (T428883)
- 09:29 marostegui@cumin1003: dbctl commit (dc=all): 'Set es1037 with weight 0 T429436', diff saved to https://phabricator.wikimedia.org/P94223 and previous config saved to /var/cache/conftool/dbconfig/20260617-092940-marostegui.json
- 09:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: Primary switchover es6 T429436
- 09:29 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs1002.eqiad.wmnet with OS bookworm
- 09:29 marostegui@cumin1003: dbctl commit (dc=all): 'Set es6 eqiad as read-only for maintenance - T429436', diff saved to https://phabricator.wikimedia.org/P94222 and previous config saved to /var/cache/conftool/dbconfig/20260617-092913-marostegui.json
- 09:27 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2203.codfw.wmnet with reason: host reimage
- 09:26 jynus: testing x1 backups @ cumin2003 T427897
- 09:11 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2203.codfw.wmnet with OS trixie
- 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2203: Upgrading db2203.codfw.wmnet
- 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2203: Upgrading db2203.codfw.wmnet
- 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 09:07 elukey: add basic Kafka ACLs for anonymous to logging-codfw - T425528 (I'll add rollback steps in the task if needed)
- 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2045: repool after maintenance es2045
- 09:06 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 09:05 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool es2044: Upgrading es2044.codfw.wmnet
- 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2044: Upgrading es2044.codfw.wmnet
- 09:04 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Promote es2046 to es5 codfw primary T428572', diff saved to https://phabricator.wikimedia.org/P94219 and previous config saved to /var/cache/conftool/dbconfig/20260617-090221-marostegui.json
- 09:02 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 09:01 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 09:00 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
- 08:59 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
- 08:57 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2203 T429190', diff saved to https://phabricator.wikimedia.org/P94218 and previous config saved to /var/cache/conftool/dbconfig/20260617-085615-cwilliams.json
- 08:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2009.codfw.wmnet with OS trixie
- 08:55 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:55 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2212 to s1 primary T429190', diff saved to https://phabricator.wikimedia.org/P94217 and previous config saved to /var/cache/conftool/dbconfig/20260617-085310-cwilliams.json
- 08:51 cezmunsta: Starting s1 codfw failover from db2203 to db2212 - T429190
- 08:51 marostegui@dns1004: END - running authdns-update
- 08:49 marostegui@dns1004: START - running authdns-update
- 08:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2212 with weight 0 T429190', diff saved to https://phabricator.wikimedia.org/P94215 and previous config saved to /var/cache/conftool/dbconfig/20260617-084642-cwilliams.json
- 08:46 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 08:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s1 T429190
- 08:45 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1044: repool after upgrade
- 08:38 jelto: "Imported helm3 3.19.5-1 to bullseye-wikimedia, bookworm-wikimedia and trixie-wikimedia - T427403"
- 08:38 moritzm: installing apache2 security updates
- 08:36 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2009.codfw.wmnet with reason: host reimage
- 08:31 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2009.codfw.wmnet with reason: host reimage
- 08:25 mlitn@deploy1003: Finished scap sync-world: Backport for Squashed diff to master, Squashed diff to master (duration: 35m 34s)
- 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2008.codfw.wmnet with OS trixie
- 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:22 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:17 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 08:14 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host conf2009.codfw.wmnet with OS trixie
- 08:12 mlitn@deploy1003: mlitn: Continuing with deployment
- 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 08:09 mlitn@deploy1003: mlitn: Backport for Squashed diff to master, Squashed diff to master synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 08:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 08:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 08:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 08:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2008.codfw.wmnet with reason: host reimage
- 08:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2007.codfw.wmnet with OS trixie
- 08:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:03 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 08:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-wdqs1001.eqiad.wmnet with OS bookworm
- 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1044: repool after upgrade
- 08:00 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2008.codfw.wmnet with reason: host reimage
- 07:59 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 07:58 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:57 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1044.eqiad.wmnet with OS trixie
- 07:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2009.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:49 mlitn@deploy1003: Started scap sync-world: Backport for Squashed diff to master, Squashed diff to master
- 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2007.codfw.wmnet with reason: host reimage
- 07:43 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 07:43 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host conf2008.codfw.wmnet with OS trixie
- 07:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:40 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1044.eqiad.wmnet with reason: host reimage
- 07:39 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2007.codfw.wmnet with reason: host reimage
- 07:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1044.eqiad.wmnet with reason: host reimage
- 07:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 07:23 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 07:23 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host conf2007.codfw.wmnet with OS trixie
- 07:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 07:22 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes (attempt 3) - oblivian@cumin1003"
- 07:22 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes (attempt 3) - oblivian@cumin1003
- 07:21 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
- 07:21 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes (attempt 3) - oblivian@cumin1003
- 07:21 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes (attempt 3) - oblivian@cumin1003"
- 07:17 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1044.eqiad.wmnet with OS trixie
- 07:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1044: Upgrading es1044.eqiad.wmnet
- 07:15 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1044: Upgrading es1044.eqiad.wmnet
- 07:15 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: Migration of es1037.eqiad.wmnet completed
- 06:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "revert deployment - oblivian@cumin1003"
- 06:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: revert deployment - oblivian@cumin1003
- 06:52 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: revert deployment - oblivian@cumin1003
- 06:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "revert deployment - oblivian@cumin1003"
- 06:46 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes - oblivian@cumin1003"
- 06:46 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes - oblivian@cumin1003
- 06:46 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes - oblivian@cumin1003
- 06:46 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes - oblivian@cumin1003"
- 06:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: Migration of es1037.eqiad.wmnet completed
- 06:16 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1037.eqiad.wmnet with OS trixie
- 05:59 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1037.eqiad.wmnet with reason: host reimage
- 05:54 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1037.eqiad.wmnet with reason: host reimage
- 05:38 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1037.eqiad.wmnet with OS trixie
- 05:37 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: Upgrading es1037.eqiad.wmnet
- 05:37 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: Upgrading es1037.eqiad.wmnet
- 05:37 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s)
- 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:01 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 00:01 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - pt1979@cumin2002"
- 00:01 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - pt1979@cumin2002"
2026-06-16
- 23:44 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage
- 23:38 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage
- 23:03 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 23:02 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-restart-haproxy (exit_code=0) rolling restart of HAProxy on A:cp - OpenSSL update ()
- 23:01 pt1979@cumin2002: END (FAIL) - Cookbook sre.hosts.dhcp (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet
- 22:57 pt1979@cumin2002: START - Cookbook sre.hosts.dhcp for host wikikube-ctrl2006.codfw.wmnet
- 22:57 pt1979@cumin2002: END (FAIL) - Cookbook sre.hosts.dhcp (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet
- 22:52 pt1979@cumin2002: START - Cookbook sre.hosts.dhcp for host wikikube-ctrl2006.codfw.wmnet
- 22:50 pt1979@cumin2002: END (FAIL) - Cookbook sre.hosts.dhcp (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet
- 22:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 22:49 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 22:37 pt1979@cumin2002: START - Cookbook sre.hosts.dhcp for host wikikube-ctrl2006.codfw.wmnet
- 22:30 robh@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS bookworm
- 22:09 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 22:08 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 22:07 kemayo@deploy1003: Finished scap sync-world: Backport for Update VE core submodule to master (0930c3a9e) (T406841 T429174 T397501 T424632 T429355), Update VE core submodule to master (0930c3a9e) (T397501 T424632 T429355) (duration: 08m 11s)
- 22:02 kemayo@deploy1003: kemayo: Continuing with deployment
- 22:01 kemayo@deploy1003: kemayo: Backport for Update VE core submodule to master (0930c3a9e) (T406841 T429174 T397501 T424632 T429355), Update VE core submodule to master (0930c3a9e) (T397501 T424632 T429355) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:59 kemayo@deploy1003: Started scap sync-world: Backport for Update VE core submodule to master (0930c3a9e) (T406841 T429174 T397501 T424632 T429355), Update VE core submodule to master (0930c3a9e) (T397501 T424632 T429355)
- 21:52 ryankemper@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 21:50 ryankemper@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 21:49 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS bookworm
- 21:48 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 21:48 robh@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 21:46 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:46 ryankemper@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 21:46 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:46 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:46 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
- 21:45 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 21:38 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 21:34 cscott@deploy1003: Finished scap sync-world: Backport for Update definition of html heading to match Parsoid/core (T417530 T417531 T428677) (duration: 18m 41s)
- 21:32 robh@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 21:31 robh@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 21:30 robh@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 21:29 cscott@deploy1003: arlolra, cscott: Continuing with deployment
- 21:26 urbanecm@deploy1003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply
- 21:25 urbanecm@deploy1003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply
- 21:24 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply
- 21:24 robh@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-ctrl2006.codfw.wmnet with OS bookworm
- 21:23 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply
- 21:21 urbanecm@deploy1003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply
- 21:20 urbanecm@deploy1003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply
- 21:20 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS bookworm
- 21:17 cscott@deploy1003: arlolra, cscott: Backport for Update definition of html heading to match Parsoid/core (T417530 T417531 T428677) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:15 cscott@deploy1003: Started scap sync-world: Backport for Update definition of html heading to match Parsoid/core (T417530 T417531 T428677)
- 21:10 robh@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 21:08 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 20:54 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2043.*
- 20:51 jdlrobson@deploy1003: Finished scap sync-world: Backport for Guard round function with a supports query (T424596), Add wprov parameter to home link (T429268) (duration: 09m 28s)
- 20:47 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
- 20:43 jdlrobson@deploy1003: jdlrobson: Backport for Guard round function with a supports query (T424596), Add wprov parameter to home link (T429268) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:41 jdlrobson@deploy1003: Started scap sync-world: Backport for Guard round function with a supports query (T424596), Add wprov parameter to home link (T429268)
- 20:40 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.*
- 20:33 brett@dns1004: END - running authdns-update
- 20:31 brett@dns1004: START - running authdns-update
- 20:30 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5004.wikimedia.org with OS bookworm
- 20:30 brett@dns5004: FAIL - running authdns-update
- 20:29 brett@dns5004: START - running authdns-update
- 20:28 brett@dns5004: FAIL - running authdns-update
- 20:27 kemayo@deploy1003: Finished scap sync-world: Backport for EditChecks: Namespace tracking object for seen/shown/used checks (duration: 09m 50s)
- 20:26 brett@dns5004: START - running authdns-update
- 20:26 brett@dns5004: START - running authdns-update
- 20:25 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.*,service=authdns-update
- 20:23 kemayo@deploy1003: kemayo: Continuing with deployment
- 20:19 kemayo@deploy1003: kemayo: Backport for EditChecks: Namespace tracking object for seen/shown/used checks synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:18 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
- 20:17 kemayo@deploy1003: Started scap sync-world: Backport for EditChecks: Namespace tracking object for seen/shown/used checks
- 20:09 jasmine@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 20:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs1001.eqiad.wmnet with reason: host reimage
- 19:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-wdqs1001.eqiad.wmnet with reason: host reimage
- 19:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 19:55 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 19:54 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:47 jasmine@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:46 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:45 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS bookworm
- 19:45 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs1001.eqiad.wmnet with OS bookworm
- 19:39 jasmine@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:35 brett@cumin2002: START - Cookbook sre.cdn.roll-restart-haproxy rolling restart of HAProxy on A:cp - OpenSSL update ()
- 19:34 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:31 jhancock@cumin2002: START - Cookbook sre.dns.netbox
- 19:30 brett@cumin2002: START - Cookbook sre.cdn.roll-restart-haproxy rolling restart of HAProxy on A:cp - OpenSSL update ()
- 19:27 jasmine@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 19:18 topranks: restarting grpc server on eqiad SR-Linux switches to recover from problem of no free threads T429242
- 19:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 19:08 robh@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 19:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:00 krinkle@deploy1003: Finished scap sync-world: Backport for Disable ShortUrl on hiwiki, hiwikiversity, maiwiki, knwiki, knwikisource, tcywiki (T107188) (duration: 11m 18s)
- 18:58 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 18:56 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 18:56 jasmine@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 18:55 krinkle@deploy1003: krinkle: Continuing with deployment
- 18:52 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 18:51 krinkle@deploy1003: krinkle: Backport for Disable ShortUrl on hiwiki, hiwikiversity, maiwiki, knwiki, knwikisource, tcywiki (T107188) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 18:48 krinkle@deploy1003: Started scap sync-world: Backport for Disable ShortUrl on hiwiki, hiwikiversity, maiwiki, knwiki, knwikisource, tcywiki (T107188)
- 18:45 jasmine@cumin2002: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 18:44 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage
- 18:41 eevans@deploy1003: helmfile [codfw] DONE helmfile.d/services/data-gateway: apply
- 18:41 eevans@deploy1003: helmfile [codfw] START helmfile.d/services/data-gateway: apply
- 18:41 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage
- 18:40 eevans@deploy1003: helmfile [eqiad] DONE helmfile.d/services/data-gateway: apply
- 18:39 eevans@deploy1003: helmfile [eqiad] START helmfile.d/services/data-gateway: apply
- 18:39 eevans@deploy1003: helmfile [staging] DONE helmfile.d/services/data-gateway: apply
- 18:39 eevans@deploy1003: helmfile [staging] START helmfile.d/services/data-gateway: apply
- 18:35 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 18:34 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 18:33 robh@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 18:30 robh@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
- 18:23 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.7 refs T423916
- 18:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 18:12 brett@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host dns5004
- 18:12 brett@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dns5004
- 18:08 brett@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host dns5004
- 18:08 brett@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dns5004.wikimedia.org 8.166.102.103.in-addr.arpa 8.0.0.0.6.6.1.0.2.0.1.0.3.0.1.0.1.0.0.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 18:08 brett@cumin2002: START - Cookbook sre.dns.wipe-cache dns5004.wikimedia.org 8.166.102.103.in-addr.arpa 8.0.0.0.6.6.1.0.2.0.1.0.3.0.1.0.1.0.0.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 18:08 brett@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 18:08 brett@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host dns5004 - brett@cumin2002"
- 18:08 brett@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host dns5004 - brett@cumin2002"
- 18:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 18:00 brett@cumin2002: START - Cookbook sre.dns.netbox
- 18:00 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 17:59 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 17:53 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.*
- 17:47 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 17:47 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: change mgmt name for frproto1001 - cmooney@cumin1003"
- 17:46 brett@cumin2002: START - Cookbook sre.hosts.move-vlan for host dns5004
- 17:46 brett@cumin2002: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS bookworm
- 17:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: change mgmt name for frproto1001 - cmooney@cumin1003"
- 17:43 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host conf2007.codfw.wmnet with OS trixie
- 17:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Revert^2 "hCaptcha: Enable for UploadWizard on all wikis with it", PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322), PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322) (duration: 32m 19s)
- 17:38 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 17:30 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 17:29 dreamyjazz@deploy1003: dreamyjazz: Backport for Revert^2 "hCaptcha: Enable for UploadWizard on all wikis with it", PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322), PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t
- 17:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host conf2007.codfw.wmnet with OS trixie
- 17:25 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host kafka-logging1007.eqiad.wmnet with OS trixie
- 17:20 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1007.eqiad.wmnet with OS trixie
- 17:11 dreamyjazz@deploy1003: Started scap sync-world: Backport for Revert^2 "hCaptcha: Enable for UploadWizard on all wikis with it", PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322), PublishCaptchaHandler: Only require CAPTCHA for UploadWizard (T429322)
- 16:35 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 16:09 brennen@deploy1003: Finished deploy [phabricator/deployment@a640ed9]: deploy phab1004 - T429350 (duration: 00m 45s)
- 16:08 brennen@deploy1003: Started deploy [phabricator/deployment@a640ed9]: deploy phab1004 - T429350
- 16:08 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab1004.eqiad.wmnet with reason: Phorge Deploy
- 16:08 brennen@deploy1003: Finished deploy [phabricator/deployment@a640ed9]: deploy phab2002 - T429350 (duration: 00m 47s)
- 16:07 brennen@deploy1003: Started deploy [phabricator/deployment@a640ed9]: deploy phab2002 - T429350
- 16:06 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2002.codfw.wmnet with reason: Phorge Deploy
- 16:04 cmooney@cumin2002: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: codfw rack a5 depool for switch maintenance T428020
- 15:42 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:migrateMentorStatusAway --wiki=abwiki --dry-run # T409170
- 15:39 moritzm: installing Tomcat security updates
- 15:38 urbanecm: Remove `migrateMentorStatusAwayToCommunityConfiguration` from `updatelog` on all wikis in `growthexperiments.dblist` (T409170)
- 15:38 dancy@deploy1003: Installation of scap version "4.269.0" completed for 2 hosts
- 15:36 dancy@deploy1003: Installing scap version "4.269.0" for 2 host(s)
- 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@a640ed9]: test deploy phab2003 - T427286 (duration: 00m 49s)
- 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@a640ed9]: test deploy phab2003 - T427286
- 15:16 cmooney@cumin2002: START - Cookbook sre.mysql.pool pool db2176: codfw rack a5 depool for switch maintenance T428020
- 15:16 cmooney@cumin2002: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: codfw rack a5 depool for switch maintenance T428020
- 15:07 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments purgeUserOptions.php --login-age 1 growthexperiments-tour-homepage-welcome # T429352
- 15:06 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments purgeUserOptions.php --login-age 1 growthexperiments-tour-homepage-discovery # T429352
- 15:03 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments purgeUserOptions.php --login-age 1 growthexperiments-tour-homepage-mentorship # T429352
- 15:01 awight@deploy1003: Finished scap sync-world: Backport for Hotfix for T428620 (T428620) (duration: 10m 00s)
- 14:57 awight@deploy1003: seanleong-wmde, awight: Continuing with deployment
- 14:55 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments purgeUserOptions.php --login-age 1 growthexperiments-tour-help-panel # T429352
- 14:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 14:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update records for frproto1001 (formerly payments1008) - cmooney@cumin1003"
- 14:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update records for frproto1001 (formerly payments1008) - cmooney@cumin1003"
- 14:53 awight@deploy1003: seanleong-wmde, awight: Backport for Hotfix for T428620 (T428620) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:51 awight@deploy1003: Started scap sync-world: Backport for Hotfix for T428620 (T428620)
- 14:48 aokoth@deploy1003: Finished deploy [phabricator/deployment@73e57ce]: deploy phab (duration: 02m 09s)
- 14:46 aokoth@deploy1003: Started deploy [phabricator/deployment@73e57ce]: deploy phab
- 14:28 cmooney@cumin2002: START - Cookbook sre.mysql.pool pool db2175: codfw rack a5 depool for switch maintenance T428020
- 14:28 cmooney@cumin2002: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: codfw rack a5 depool for switch maintenance T428020
- 14:07 dcausse@deploy1003: Finished scap sync-world: Backport for Bump wikimedia/parsoid to 0.24.0-a10 (T417530 T428105 T429187), Bump wikimedia/parsoid to 0.24.0-a10 (T429187) (duration: 11m 29s)
- 14:07 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 14:03 dcausse@deploy1003: jgiannelos, dcausse: Continuing with deployment
- 14:02 cmooney@cumin1003: START - Cookbook sre.dns.netbox
- 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:59 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:58 dcausse@deploy1003: jgiannelos, dcausse: Backport for Bump wikimedia/parsoid to 0.24.0-a10 (T417530 T428105 T429187), Bump wikimedia/parsoid to 0.24.0-a10 (T429187) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:56 dcausse@deploy1003: Started scap sync-world: Backport for Bump wikimedia/parsoid to 0.24.0-a10 (T417530 T428105 T429187), Bump wikimedia/parsoid to 0.24.0-a10 (T429187)
- 13:54 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 13:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 13:52 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 13:52 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 13:51 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 13:48 atsuko@deploy1003: Finished scap sync-world: Backport for Revert "translate: remove CirrusSearch endpoints" (duration: 04m 10s)
- 13:47 atsuko@deploy1003: atsuko: Rolling back deployment
- 13:47 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 13:46 atsuko@deploy1003: atsuko: Backport for Revert "translate: remove CirrusSearch endpoints" synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:44 atsuko@deploy1003: Started scap sync-world: Backport for Revert "translate: remove CirrusSearch endpoints"
- 13:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 13:43 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 13:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 13:40 cmooney@cumin2002: START - Cookbook sre.mysql.pool pool db2157: codfw rack a5 depool for switch maintenance T428020
- 13:40 cmooney@cumin2002: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: codfw rack a5 depool for switch maintenance T428020
- 13:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 13:39 atsuko@deploy1003: Finished scap sync-world: Backport for translate: remove CirrusSearch endpoints (T425377) (duration: 11m 16s)
- 13:37 atsuko@deploy1003: atsuko: Rolling back deployment
- 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1080.eqiad.wmnet with OS trixie
- 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:36 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:34 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: codfw rack a5 depool for switch maintenance T428020
- 13:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1079.eqiad.wmnet with OS trixie
- 13:32 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:30 atsuko@deploy1003: atsuko: Backport for translate: remove CirrusSearch endpoints (T425377) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:28 atsuko@deploy1003: Started scap sync-world: Backport for translate: remove CirrusSearch endpoints (T425377)
- 13:25 dcausse@deploy1003: Finished scap sync-world: Backport for Replace wgNewUserMessageOnAutoCreate with wgNewUserMessageOnFirstEdit (T426206) (duration: 08m 50s)
- 13:25 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:21 dcausse@deploy1003: dcausse, neriah: Continuing with deployment
- 13:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
- 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1080.eqiad.wmnet with reason: host reimage
- 13:18 dcausse@deploy1003: dcausse, neriah: Backport for Replace wgNewUserMessageOnAutoCreate with wgNewUserMessageOnFirstEdit (T426206) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:16 dcausse@deploy1003: Started scap sync-world: Backport for Replace wgNewUserMessageOnAutoCreate with wgNewUserMessageOnFirstEdit (T426206)
- 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:12 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1080.eqiad.wmnet with reason: host reimage
- 13:12 mfossati@deploy1003: Finished scap sync-world: Backport for Remove custom streams (T423148) (duration: 08m 35s)
- 13:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1079.eqiad.wmnet with reason: host reimage
- 13:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging1008.eqiad.wmnet with OS trixie
- 13:08 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:07 jmm@dns1004: END - running authdns-update
- 13:06 mfossati@deploy1003: ksarabia, mfossati: Continuing with deployment
- 13:05 mfossati@deploy1003: ksarabia, mfossati: Backport for Remove custom streams (T423148) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:05 jmm@dns1004: START - running authdns-update
- 13:03 mfossati@deploy1003: Started scap sync-world: Backport for Remove custom streams (T423148)
- 13:02 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1079.eqiad.wmnet with reason: host reimage
- 13:02 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 13:02 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1080.eqiad.wmnet with OS trixie
- 12:57 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:52 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1079.eqiad.wmnet with OS trixie
- 12:52 cmooney@cumin2002: START - Cookbook sre.mysql.pool pool db2154: codfw rack a5 depool for switch maintenance T428020
- 12:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host kafka-logging1007.eqiad.wmnet with OS trixie
- 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-logging1006.eqiad.wmnet with OS trixie
- 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 12:49 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
- 12:48 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2153: codfw rack a5 depool for switch maintenance T428020
- 12:47 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2255.codfw.wmnet
- 12:47 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2255.codfw.wmnet
- 12:47 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2254.codfw.wmnet
- 12:47 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2254.codfw.wmnet
- 12:47 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2243.codfw.wmnet
- 12:47 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2243.codfw.wmnet
- 12:47 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2242.codfw.wmnet
- 12:47 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2242.codfw.wmnet
- 12:47 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 12:47 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2092.codfw.wmnet
- 12:47 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2092.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2091.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2091.codfw.wmnet
- 12:46 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2078.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2078.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2077.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2077.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2076.codfw.wmnet
- 12:46 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2076.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2075.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2075.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2074.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2074.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2051.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2051.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2044.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2044.codfw.wmnet
- 12:46 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2041.codfw.wmnet
- 12:46 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2041.codfw.wmnet
- 12:46 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet
- 12:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:45 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet
- 12:45 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2018.codfw.wmnet
- 12:45 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2018.codfw.wmnet
- 12:45 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2017.codfw.wmnet
- 12:45 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2017.codfw.wmnet
- 12:45 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2014.codfw.wmnet
- 12:45 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2014.codfw.wmnet
- 12:45 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2013.codfw.wmnet
- 12:45 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2013.codfw.wmnet
- 12:45 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2012.codfw.wmnet
- 12:45 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2012.codfw.wmnet
- 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging1008.eqiad.wmnet with reason: host reimage
- 12:43 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
- 12:40 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging1008.eqiad.wmnet with reason: host reimage
- 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-logging1006.eqiad.wmnet with reason: host reimage
- 12:28 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1008.eqiad.wmnet with OS trixie
- 12:24 topranks: reboot lsw1-a5-codfw to complete JunOS upgrade T428020
- 12:23 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1007.eqiad.wmnet with OS trixie
- 12:22 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-logging1006.eqiad.wmnet with reason: host reimage
- 12:19 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2255.codfw.wmnet
- 12:19 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2255.codfw.wmnet
- 12:19 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2254.codfw.wmnet
- 12:18 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2254.codfw.wmnet
- 12:17 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2243.codfw.wmnet
- 12:17 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2243.codfw.wmnet
- 12:17 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2242.codfw.wmnet
- 12:16 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2242.codfw.wmnet
- 12:16 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2092.codfw.wmnet
- 12:16 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2092.codfw.wmnet
- 12:16 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2091.codfw.wmnet
- 12:15 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2091.codfw.wmnet
- 12:15 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2078.codfw.wmnet
- 12:14 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2078.codfw.wmnet
- 12:14 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2077.codfw.wmnet
- 12:14 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2077.codfw.wmnet
- 12:14 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2076.codfw.wmnet
- 12:13 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2076.codfw.wmnet
- 12:13 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2075.codfw.wmnet
- 12:12 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2075.codfw.wmnet
- 12:12 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2074.codfw.wmnet
- 12:12 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2074.codfw.wmnet
- 12:12 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2051.codfw.wmnet
- 12:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-a5-codfw JunOS upgrade
- 12:07 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2051.codfw.wmnet
- 12:06 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a5-codfw,lsw1-a5-codfw IPv6,lsw1-a5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgrrade
- 12:06 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2044.codfw.wmnet
- 12:06 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2044.codfw.wmnet
- 12:06 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2041.codfw.wmnet
- 12:05 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2041.codfw.wmnet
- 12:05 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2018.codfw.wmnet
- 12:05 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2018.codfw.wmnet
- 12:04 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2017.codfw.wmnet
- 12:04 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2017.codfw.wmnet
- 12:04 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2014.codfw.wmnet
- 12:03 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2014.codfw.wmnet
- 12:03 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2013.codfw.wmnet
- 12:03 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2013.codfw.wmnet
- 12:02 cmooney@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2012.codfw.wmnet
- 12:02 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet
- 12:01 cmooney@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2012.codfw.wmnet
- 12:01 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1006.eqiad.wmnet with OS trixie
- 11:57 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet
- 11:51 dreamyjazz@deploy1003: Finished scap sync-world: Backport for Revert "hCaptcha: Enable for UploadWizard on all wikis with it" (duration: 08m 45s)
- 11:49 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2176: codfw rack a5 depool for switch maintenance T428020
- 11:49 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2176: codfw rack a5 depool for switch maintenance T428020
- 11:49 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2175: codfw rack a5 depool for switch maintenance T428020
- 11:48 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2175: codfw rack a5 depool for switch maintenance T428020
- 11:48 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2157: codfw rack a5 depool for switch maintenance T428020
- 11:48 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1078
- 11:48 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2157: codfw rack a5 depool for switch maintenance T428020
- 11:48 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2154: codfw rack a5 depool for switch maintenance T428020
- 11:47 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2154: codfw rack a5 depool for switch maintenance T428020
- 11:47 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 11:46 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2153: codfw rack a5 depool for switch maintenance T428020
- 11:46 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2153: codfw rack a5 depool for switch maintenance T428020
- 11:46 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1078
- 11:46 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:45 dreamyjazz@deploy1003: dreamyjazz: Backport for Revert "hCaptcha: Enable for UploadWizard on all wikis with it" synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:43 jclark@cumin1003: START - Cookbook sre.dns.netbox
- 11:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for Revert "hCaptcha: Enable for UploadWizard on all wikis with it"
- 11:42 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1078
- 11:41 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1078
- 11:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 11:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Migration of es2035.codfw.wmnet completed
- 11:06 moritzm: installing Bird security updates on routed Ganeti nodes
- 10:49 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es1037 T429118', diff saved to https://phabricator.wikimedia.org/P94172 and previous config saved to /var/cache/conftool/dbconfig/20260616-104931-marostegui.json
- 10:25 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
- 10:24 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
- 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Migration of es2035.codfw.wmnet completed
- 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-redacteddb1001.eqiad.wmnet
- 10:24 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-redacteddb1001.eqiad.wmnet
- 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts
- 10:24 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts
- 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet
- 10:24 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet
- 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1154.eqiad.wmnet
- 10:24 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1154.eqiad.wmnet
- 10:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 10:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Migration of es1036.eqiad.wmnet completed
- 10:22 jmm@dns1004: END - running authdns-update
- 10:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 10:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 10:21 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 10:21 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 10:20 jmm@dns1004: START - running authdns-update
- 10:20 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 10:19 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
- 10:18 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
- 10:18 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:18 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
- 10:18 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
- 10:17 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 10:17 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2035.codfw.wmnet with OS trixie
- 09:59 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2035.codfw.wmnet with reason: host reimage
- 09:52 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2035.codfw.wmnet with reason: host reimage
- 09:49 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
- 09:48 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
- 09:47 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for UploadWizard on all wikis with it (T426126) (duration: 09m 38s)
- 09:43 marostegui: Drop wrongly created table son testwikidatawiki s3 master T429304
- 09:42 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 09:39 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for UploadWizard on all wikis with it (T426126) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:38 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/refreshUserImpactData.php --wiki=wikidatawiki --registeredWithin=2week --hasEditsAtLeast=3 --ignoreIfUpdatedWithin=6hour --verbose --use-job-queue # T418115
- 09:37 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for UploadWizard on all wikis with it (T426126)
- 09:37 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/refreshUserImpactData.php --wiki=wikidatawiki --registeredWithin=1year --editedWithin=2week --hasEditsAtLeast=3 --ignoreIfUpdatedWithin=6hour --verbose --use-job-queue # T418115
- 09:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Migration of es1036.eqiad.wmnet completed
- 09:37 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/refreshUserImpactData.php --registeredWithin=2week --hasEditsAtLeast=3 --ignoreIfUpdatedWithin=6hour --verbose --use-job-queue # T418115
- 09:35 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2035.codfw.wmnet with OS trixie
- 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2035: Upgrading es2035.codfw.wmnet
- 09:34 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2035: Upgrading es2035.codfw.wmnet
- 09:34 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 09:32 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es2035 T429303', diff saved to https://phabricator.wikimedia.org/P94164 and previous config saved to /var/cache/conftool/dbconfig/20260616-093247-marostegui.json
- 09:31 marostegui@cumin1003: dbctl commit (dc=all): 'Promote es2037 to es6 primary T429303', diff saved to https://phabricator.wikimedia.org/P94163 and previous config saved to /var/cache/conftool/dbconfig/20260616-093149-marostegui.json
- 09:31 jayme: imported istioctl 1.29.4-1 to bookworm-/trixie-wikimedia - T427401
- 09:30 marostegui: Starting es6 codfw failover from es2035 to es2037 - T429303
- 09:30 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 09:30 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 09:30 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 09:29 marostegui@cumin1003: dbctl commit (dc=all): 'Set es2037 with weight 0 T429303', diff saved to https://phabricator.wikimedia.org/P94162 and previous config saved to /var/cache/conftool/dbconfig/20260616-092937-marostegui.json
- 09:29 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 09:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: Primary switchover es6 T429303
- 09:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1036.eqiad.wmnet with OS trixie
- 09:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:24 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:23 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:21 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:20 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [Growth] wikidatawiki: Enable Growth features (T418115) (duration: 16m 29s)
- 09:18 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:14 urbanecm@deploy1003: urbanecm: Continuing with deployment
- 09:13 urbanecm: php multiversion/MWScript.php WikimediaMaintenance:createExtensionTables.php --wiki={testwikidatawiki,wikidatawiki} growthexperiments # T418115, within mw-debug
- 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1036.eqiad.wmnet with reason: host reimage
- 09:07 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 09:07 tappof@cumin1003: END (PASS) - Cookbook sre.metamonitoring.downtime (exit_code=0) Downtime for 0:05:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: cookbook test
- 09:07 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 09:07 tappof@cumin1003: START - Cookbook sre.metamonitoring.downtime Downtime for 0:05:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: cookbook test
- 09:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 09:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 09:04 urbanecm@deploy1003: urbanecm: Backport for [Growth] wikidatawiki: Enable Growth features (T418115) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:04 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1036.eqiad.wmnet with reason: host reimage
- 09:02 urbanecm@deploy1003: Started scap sync-world: Backport for [Growth] wikidatawiki: Enable Growth features (T418115)
- 09:01 moritzm: uploaded bird 2.18.2-1~wmf13u1 to trixie-wikimedia T429285
- 09:00 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidata WikimediaMaintenance:createExtensionTables.php GrowthExperiments # T418115
- 08:56 moritzm: uploaded bird 2.18.2-1~wmf12u1 to bookworm-wikimedia T429285
- 08:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1036.eqiad.wmnet with OS trixie
- 08:47 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1036: Upgrading es1036.eqiad.wmnet
- 08:46 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for MobileFrontend in all wikis (T425940) (duration: 19m 23s)
- 08:45 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1036: Upgrading es1036.eqiad.wmnet
- 08:45 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: repool after upgrade
- 08:42 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 08:32 moritzm: installing nginx security updates
- 08:29 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for MobileFrontend in all wikis (T425940) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 08:27 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for MobileFrontend in all wikis (T425940)
- 08:23 mszwarc@deploy1003: Synchronized private/PrivateSettings.php: Private code deployment for Suggested Investigations (duration: 02m 23s)
- 08:19 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals: Private code deployment for Suggested Investigations (duration: 06m 03s)
- 08:17 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver codfw-k8s # T425377: populating translation memory (ttmserver-export.php) on codfw-k8s (dblist: https://phabricator.wikimedia.org/P94157)
- 08:05 wmde-fisch@deploy1003: Finished scap sync-world: Backport for Improve click intent event logging and exposure tracking (duration: 11m 31s)
- 08:00 moritzm: update bird on ganeti7001 to 2.18.2-1~wmf12u1
- 07:58 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment
- 07:58 wmde-fisch@deploy1003: wmde-fisch: Backport for Improve click intent event logging and exposure tracking synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:58 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1047: repool after upgrade
- 07:54 wmde-fisch@deploy1003: Started scap sync-world: Backport for Improve click intent event logging and exposure tracking
- 07:50 wmde-fisch@deploy1003: Finished scap sync-world: Backport for Update VE core submodule to master (3e79e9934) (T397319 T428764) (duration: 36m 13s)
- 07:36 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment
- 07:33 wmde-fisch@deploy1003: wmde-fisch: Backport for Update VE core submodule to master (3e79e9934) (T397319 T428764) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:14 wmde-fisch@deploy1003: Started scap sync-world: Backport for Update VE core submodule to master (3e79e9934) (T397319 T428764)
- 07:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1047.eqiad.wmnet with OS trixie
- 06:50 hashar@deploy1003: Finished deploy [integration/docroot@2165507]: build: Updating js-yaml to 4.2.0 (duration: 00m 16s)
- 06:50 hashar@deploy1003: Started deploy [integration/docroot@2165507]: build: Updating js-yaml to 4.2.0
- 06:44 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1047.eqiad.wmnet with reason: host reimage
- 06:40 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1047.eqiad.wmnet with reason: host reimage
- 06:25 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1047.eqiad.wmnet with OS trixie
- 06:24 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 06:24 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 06:24 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 06:24 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool es1047: Upgrading es1047.eqiad.wmnet
- 05:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1047: Upgrading es1047.eqiad.wmnet
- 05:58 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 04:55 ryankemper: T427951 Deleted 4 leftover mirrored dev/test topics from kafka-test: `eqiad.mediawiki.{page_html_content_change.dev{1,4},page_edit_type_simple.dev0}`, `eqiad.mw_page_edit_type_enrich.error`
- 04:05 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.4 (duration: 05m 29s)
2026-06-15
- 22:35 sbassett: Deployed private config for T429244
- 22:05 sbassett: Deployed updated security fix for T427611
- 22:04 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 22:04 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 22:04 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 22:03 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 21:54 dancy@deploy1003: Finished scap sync-world: Backport for beta: Point remaining db11 references at deployment-db15 (T428930) (duration: 12m 27s)
- 21:53 dancy@deploy1003: dancy: Continuing with deployment
- 21:49 dancy@deploy1003: dancy: Backport for beta: Point remaining db11 references at deployment-db15 (T428930) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:48 sbassett: Deployed security fix for T428809
- 21:48 dancy@deploy1003: Started scap sync-world: Backport for beta: Point remaining db11 references at deployment-db15 (T428930)
- 21:40 sbassett: Deployed security fix for T428820
- 21:22 sbassett@deploy1003: Finished scap sync-world: Backport for ForceReauth: Avoid unnecessary securitySensitiveOperationStatus checks (duration: 08m 11s)
- 21:17 sbassett@deploy1003: sbassett: Continuing with deployment
- 21:15 sbassett@deploy1003: sbassett: Backport for ForceReauth: Avoid unnecessary securitySensitiveOperationStatus checks synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:13 sbassett@deploy1003: Started scap sync-world: Backport for ForceReauth: Avoid unnecessary securitySensitiveOperationStatus checks
- 21:06 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp5028.*
- 21:06 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs5005.eqsin.wmnet} and A:liberica
- 21:05 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart P{lvs5005.eqsin.wmnet} and A:liberica
- 20:52 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5028.eqsin.wmnet with OS trixie
- 20:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5028.eqsin.wmnet with reason: host reimage
- 20:21 dancy@deploy1003: Finished scap sync-world: Backport for REST: set new RestModuleOverrides variable (T422756), Enable "exit the editor" survey on 11 wikis for phase 2 (T426132) (duration: 10m 54s)
- 20:17 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5028.eqsin.wmnet with reason: host reimage
- 20:16 dancy@deploy1003: caro, dancy, bpirkle: Continuing with deployment
- 20:14 dancy@deploy1003: caro, dancy, bpirkle: Backport for REST: set new RestModuleOverrides variable (T422756), Enable "exit the editor" survey on 11 wikis for phase 2 (T426132) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:10 dancy@deploy1003: Started scap sync-world: Backport for REST: set new RestModuleOverrides variable (T422756), Enable "exit the editor" survey on 11 wikis for phase 2 (T426132)
- 20:02 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dse-k8s-wdqs2001.codfw.wmnet with OS trixie
- 19:44 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host dse-k8s-wdqs2001.codfw.wmnet with OS trixie
- 19:44 brett@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5028
- 19:44 brett@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp5028
- 19:43 brett@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host cp5028
- 19:43 brett@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5028.eqsin.wmnet 25.0.132.10.in-addr.arpa 5.2.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 19:43 brett@cumin2002: START - Cookbook sre.dns.wipe-cache cp5028.eqsin.wmnet 25.0.132.10.in-addr.arpa 5.2.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 19:43 brett@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 19:43 brett@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cp5028 - brett@cumin2002"
- 19:42 brett@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cp5028 - brett@cumin2002"
- 19:40 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host dse-k8s-wdqs2001.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:36 brett@cumin2002: START - Cookbook sre.dns.netbox
- 19:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3067.esams.wmnet
- 19:34 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3067.esams.wmnet
- 19:33 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp5026.*
- 19:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3066.esams.wmnet
- 19:33 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host dse-k8s-wdqs2001.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 19:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3066.esams.wmnet
- 19:26 brett@cumin2002: START - Cookbook sre.hosts.move-vlan for host cp5028
- 19:25 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp5028.eqsin.wmnet with OS trixie
- 19:23 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin
- 19:21 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin
- 19:18 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp5026.*
- 19:17 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs5005.eqsin.wmnet} and A:liberica
- 19:16 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart P{lvs5005.eqsin.wmnet} and A:liberica
- 19:15 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs5004.eqsin.wmnet} and A:liberica
- 19:14 brett@cumin2002: START - Cookbook sre.loadbalancer.admin config_reloading P{lvs5004.eqsin.wmnet} and A:liberica
- 19:06 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp5026.*
- 19:05 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp5026.*
- 19:05 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs5005.eqsin.wmnet} and A:liberica
- 19:04 brett@cumin2002: START - Cookbook sre.loadbalancer.admin config_reloading P{lvs5005.eqsin.wmnet} and A:liberica
- 19:04 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
- 18:44 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on P{cp7001.magru.wmnet} and A:cp
- 18:42 brett@cumin2002: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on P{cp7001.magru.wmnet} and A:cp
- 18:35 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
- 18:27 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs5005.eqsin.wmnet} and A:liberica
- 18:27 brett@cumin2002: START - Cookbook sre.loadbalancer.admin config_reloading P{lvs5005.eqsin.wmnet} and A:liberica
- 18:27 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
- 18:18 mutante: releases2003 - systemctl stop tmp.mount
- 17:53 brett@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5026
- 17:53 brett@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp5026
- 17:52 brett@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host cp5026
- 17:52 brett@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5026.eqsin.wmnet 37.0.132.10.in-addr.arpa 7.3.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 17:52 brett@cumin2002: START - Cookbook sre.dns.wipe-cache cp5026.eqsin.wmnet 37.0.132.10.in-addr.arpa 7.3.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 17:52 brett@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 17:52 brett@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cp5026 - brett@cumin2002"
- 17:52 brett@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cp5026 - brett@cumin2002"
- 17:46 brett@cumin2002: START - Cookbook sre.dns.netbox
- 17:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
- 17:40 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
- 17:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad
- 17:35 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad
- 17:34 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad
- 17:34 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad
- 17:09 brett@cumin2002: START - Cookbook sre.hosts.move-vlan for host cp5026
- 17:07 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
- 17:03 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 16:36 atsuko@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
- 16:36 atsuko@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
- 16:16 atsuko@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
- 16:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 16:16 atsuko@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
- {{safesubst:SAL entry|1=16:13 dreamyjazz@deploy1003: Finished scap sync-world: Backport for SourceEditorOverlayHookPayload: Allow aborting of the save (T428287), hCaptcha MobileFrontend: Avoid indefinite save loop on known errors (T428287), OATHUserRepository: Specify caller in query, Bump guzzlehttp/psr to version 2.11.0 (T429208), [[gerrit:1302169|NoReferrerLinks: Add re}}
- 16:13 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 16:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 16:08 atsuko@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
- 16:08 dreamyjazz@deploy1003: reedy, dreamyjazz, kharlan: Continuing with deployment
- 16:08 atsuko@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
- {{safesubst:SAL entry|1=16:07 dreamyjazz@deploy1003: reedy, dreamyjazz, kharlan: Backport for SourceEditorOverlayHookPayload: Allow aborting of the save (T428287), hCaptcha MobileFrontend: Avoid indefinite save loop on known errors (T428287), OATHUserRepository: Specify caller in query, Bump guzzlehttp/psr to version 2.11.0 (T429208), [[gerrit:1302169|NoReferrerLinks: Add}}
- {{safesubst:SAL entry|1=16:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for SourceEditorOverlayHookPayload: Allow aborting of the save (T428287), hCaptcha MobileFrontend: Avoid indefinite save loop on known errors (T428287), OATHUserRepository: Specify caller in query, Bump guzzlehttp/psr to version 2.11.0 (T429208), [[gerrit:1302169|NoReferrerLinks: Add rel}}
- 16:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 16:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:51 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:51 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet with reason: puppet debugging
- 15:50 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases1003.eqiad.wmnet with reason: puppet debugging
- 15:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2008.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 15:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Migration of db1196.eqiad.wmnet completed
- 15:41 mutante: added new project language 'nyn' - Bantu language spoken by the Nkore and Hema peoples of Southwestern Uganda
- 15:40 dzahn@dns1006: END - running authdns-update
- 15:36 dzahn@dns1006: START - running authdns-update
- 15:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet
- 15:19 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet
- 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1154.eqiad.wmnet
- 15:18 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1154.eqiad.wmnet
- 15:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts
- 15:18 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts
- 15:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-redacteddb1001.eqiad.wmnet
- 15:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-redacteddb1001.eqiad.wmnet
- 15:16 topranks: repool esams following cr2-esams rpd crash
- 15:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: no reason specified, no task ID specified]
- 15:13 cmooney@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: no reason specified, no task ID specified]
- 15:02 topranks: depool esams due to cr2-esams rpd crash
- 15:02 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: no reason specified, no task ID specified]
- 15:01 cmooney@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: no reason specified, no task ID specified]
- 15:00 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 14:58 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 14:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 14:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Migration of db1196.eqiad.wmnet completed
- 14:54 topranks: enable BGP graceful-shutdown sender on cr2-esams to drain traffic T427056
- 14:52 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-esams,cr2-esams IPv6 with reason: bouncing pic0 to reconfigure port speeds
- 14:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1196.eqiad.wmnet with OS trixie
- 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1077.eqiad.wmnet with OS trixie
- 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 14:24 elukey@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2001.codfw.wmnet with reason: tesT
- 14:24 elukey@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - elukey@cumin1003"
- 14:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1196.eqiad.wmnet with reason: host reimage
- 14:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1196.eqiad.wmnet with reason: host reimage
- 14:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 14:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 14:07 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cloudvirt1077.eqiad.wmnet with reason: host reimage
- 14:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1077.eqiad.wmnet with reason: host reimage
- 14:06 hnowlan@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 14:05 hnowlan@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 14:05 hnowlan@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 14:04 hnowlan@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 14:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1196.eqiad.wmnet with OS trixie
- 14:02 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "revert deployment - oblivian@cumin1003"
- 14:02 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: revert deployment - oblivian@cumin1003
- 14:01 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: revert deployment - oblivian@cumin1003
- 14:01 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "revert deployment - oblivian@cumin1003"
- 14:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1196: Upgrading db1196.eqiad.wmnet
- 14:00 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1196: Upgrading db1196.eqiad.wmnet
- 14:00 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 13:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1077.eqiad.wmnet with OS trixie
- 13:56 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host kafka-logging1006.eqiad.wmnet with OS trixie
- 13:54 federico3: doing a quick restart of sanitarium hosts db1155 and db1154
- 13:53 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver codfw-k8s # T425377: populating translation memory (ttmserver-export.php) on codfw-k8s (dblist: https://phabricator.wikimedia.org/P94145)
- 13:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1154.eqiad.wmnet with reason: Reboots T426633
- 13:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1155.eqiad.wmnet with reason: Reboots T426633
- 13:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 11 hosts with reason: Reboots T426633
- 13:49 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 13:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on an-redacteddb1001.eqiad.wmnet with reason: Reboots T426633
- {{safesubst:SAL entry|1=13:43 jforrester@deploy1003: Finished scap sync-world: Backport for Remove no longer used product_metrics.homepage_module_interaction (T365889 T426742), TaskSuggester: avoid nullable logger in setLogger call, migrateMentorStatusAway: ensure validateStrictly receives objects (T409170), [[gerrit:1301451|Store nowiki source in StripState::extra to support subst-nowiki (T}}
- 13:42 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 13:40 elukey@cumin1003: START - Cookbook sre.hosts.provision for host conf2007.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
- 13:39 jforrester@deploy1003: arlolra, sgimeno, jforrester: Continuing with deployment
- {{safesubst:SAL entry|1=13:37 jforrester@deploy1003: arlolra, sgimeno, jforrester: Backport for Remove no longer used product_metrics.homepage_module_interaction (T365889 T426742), TaskSuggester: avoid nullable logger in setLogger call, migrateMentorStatusAway: ensure validateStrictly receives objects (T409170), [[gerrit:1301451|Store nowiki source in StripState::extra to support subst-nowik}}
- {{safesubst:SAL entry|1=13:35 jforrester@deploy1003: Started scap sync-world: Backport for Remove no longer used product_metrics.homepage_module_interaction (T365889 T426742), TaskSuggester: avoid nullable logger in setLogger call, migrateMentorStatusAway: ensure validateStrictly receives objects (T409170), [[gerrit:1301451|Store nowiki source in StripState::extra to support subst-nowiki (T3}}
- 13:34 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host kafka-logging1006.eqiad.wmnet with OS trixie
- 13:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 13:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Migration of db2216.codfw.wmnet completed
- 13:29 topranks: enable BGP graceful-shutdown sender on cr2-esams to drain traffic T427056
- 13:28 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-esams,cr2-esams IPv6 with reason: bouncing pic0 to reconfigure port speeds
- 13:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1080.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:26 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes - oblivian@cumin1003"
- 13:25 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes - oblivian@cumin1003
- 13:25 topranks: cr2-esams, reconfigure chassis fpc to set port 0 to 100G T427056
- 13:25 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy provenance maps in HP; UX changes - oblivian@cumin1003
- 13:24 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy provenance maps in HP; UX changes - oblivian@cumin1003"
- 13:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 13:23 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Migration of db1251.eqiad.wmnet completed
- {{safesubst:SAL entry|1=13:22 jforrester@deploy1003: Finished scap sync-world: Backport for Configure wgOAuthAutoApprove['protocols'] (T412542 T426614), jawiki: remove four rights from the eliminator group (T428942), Deploy PRV to 6 wikis (T429038), [abstractwiki] Set wgForceUIMsgAsContentMsg for sidebar messages (T427730), [[gerrit:1300872|abstractwiki: Temporary config f}}
- 13:20 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 13:18 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1080.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:17 jforrester@deploy1003: arlolra, matmarex, jforrester, dragoniez: Continuing with deployment
- {{safesubst:SAL entry|1=13:13 jforrester@deploy1003: arlolra, matmarex, jforrester, dragoniez: Backport for Configure wgOAuthAutoApprove['protocols'] (T412542 T426614), jawiki: remove four rights from the eliminator group (T428942), Deploy PRV to 6 wikis (T429038), [abstractwiki] Set wgForceUIMsgAsContentMsg for sidebar messages (T427730), [[gerrit:1300872|abstractwiki: Te}}
- 13:13 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:12 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- {{safesubst:SAL entry|1=13:12 jforrester@deploy1003: Started scap sync-world: Backport for Configure wgOAuthAutoApprove['protocols'] (T412542 T426614), jawiki: remove four rights from the eliminator group (T428942), Deploy PRV to 6 wikis (T429038), [abstractwiki] Set wgForceUIMsgAsContentMsg for sidebar messages (T427730), [[gerrit:1300872|abstractwiki: Temporary config fo}}
- 13:10 moritzm: installing Linux 6.1.174 on Bookworm hosts
- 13:10 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 13:08 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 13:08 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 13:05 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 13:05 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 13:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:48 moritzm: installing augeas security updates
- 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Migration of db2216.codfw.wmnet completed
- 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging1008.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Migration of es2036.codfw.wmnet completed
- 12:38 mszwarc@deploy1003: Finished scap sync-world: Backport for Extract a service that initiates SI signal matching (T428557), Trigger Suggested Investigations when client hints are saved (T428557) (duration: 07m 42s)
- 12:37 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Migration of db1251.eqiad.wmnet completed
- 12:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2216.codfw.wmnet with OS trixie
- 12:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:34 mszwarc@deploy1003: mszwarc: Continuing with deployment
- 12:32 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1008.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:32 mszwarc@deploy1003: mszwarc: Backport for Extract a service that initiates SI signal matching (T428557), Trigger Suggested Investigations when client hints are saved (T428557) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 12:31 mszwarc@deploy1003: Started scap sync-world: Backport for Extract a service that initiates SI signal matching (T428557), Trigger Suggested Investigations when client hints are saved (T428557)
- 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1251.eqiad.wmnet with OS trixie
- 12:23 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
- 12:21 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
- 12:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2216.codfw.wmnet with reason: host reimage
- 12:15 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:12 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2216.codfw.wmnet with reason: host reimage
- 12:10 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
- 12:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1251.eqiad.wmnet with reason: host reimage
- 12:06 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
- 12:06 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host kafka-logging1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 12:02 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1251.eqiad.wmnet with reason: host reimage
- 11:56 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
- 11:55 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
- 11:54 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Migration of es2036.codfw.wmnet completed
- 11:54 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 11:53 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2216.codfw.wmnet with OS trixie
- 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2216: Upgrading db2216.codfw.wmnet
- 11:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2216: Upgrading db2216.codfw.wmnet
- 11:49 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1251.eqiad.wmnet with OS trixie
- 11:46 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1251: Upgrading db1251.eqiad.wmnet
- 11:45 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1251: Upgrading db1251.eqiad.wmnet
- 11:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 11:44 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver codfw-k8s # T425377: populating translation memory (ttmserver-export.php) on codfw-k8s (dblist: https://phabricator.wikimedia.org/P94128)
- 11:43 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host kafka-logging1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 11:43 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-k8s # T425377: populating translation memory (ttmserver-export.php) on eqiad-k8s (dblist: https://phabricator.wikimedia.org/P94127)
- 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2036.codfw.wmnet with OS trixie
- 11:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host kafka-logging1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 11:24 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2036.codfw.wmnet with reason: host reimage
- 11:17 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2036.codfw.wmnet with reason: host reimage
- 11:09 jmm@cumin2002: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling restart_daemons on A:schema-eqiad
- 11:08 jmm@cumin2002: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling restart_daemons on A:schema-eqiad
- 11:00 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2036.codfw.wmnet with OS trixie
- 10:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2036: Upgrading es2036.codfw.wmnet
- 10:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2036: Upgrading es2036.codfw.wmnet
- 10:58 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:55 jmm@cumin2002: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling restart_daemons on A:schema-codfw
- 10:54 jmm@cumin2002: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling restart_daemons on A:schema-codfw
- 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2037: repool after upgrade
- 10:52 moritzm: installing openssl security updates on bookworm
- 10:30 cgoubert@deploy1003: Finished scap sync-world: Backport for Close API Portal wiki (T427537) (duration: 07m 16s)
- 10:26 cgoubert@deploy1003: cgoubert: Continuing with deployment
- 10:25 cgoubert@deploy1003: cgoubert: Backport for Close API Portal wiki (T427537) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:23 cgoubert@deploy1003: Started scap sync-world: Backport for Close API Portal wiki (T427537)
- 10:16 blake@deploy1003: Finished scap sync-world: apache config change (T428772) (duration: 06m 41s)
- 10:12 blake@deploy1003: blake: Continuing with deployment
- 10:11 blake@deploy1003: blake: apache config change (T428772) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:10 blake@deploy1003: Started scap sync-world: apache config change (T428772)
- 10:08 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2037: repool after upgrade
- 10:04 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2037.codfw.wmnet with OS trixie
- 09:54 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:46 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 09:45 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 09:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 09:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 09:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 09:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 09:40 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-k8s # T425377: populating translation memory (ttmserver-export.php) on eqiad-k8s (dblist: https://phabricator.wikimedia.org/P94120)
- 09:35 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2037.codfw.wmnet with reason: host reimage
- 09:32 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:30 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2037.codfw.wmnet with reason: host reimage
- 09:22 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:22 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:17 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:15 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:14 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2037.codfw.wmnet with OS trixie
- 09:13 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:13 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 09:12 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:12 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 09:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 08:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 08:56 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host es2037.codfw.wmnet with OS trixie
- 08:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2037.codfw.wmnet with OS trixie
- 08:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2037: Upgrading es2037.codfw.wmnet
- 08:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2037: Upgrading es2037.codfw.wmnet
- 08:53 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 08:46 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
- 08:46 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
- 08:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
- 08:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
- 08:44 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 08:43 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 08:41 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 08:40 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 08:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 08:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 08:23 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
- 08:14 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on dbstore1008.eqiad.wmnet with reason: Maintenance
- 08:14 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1163 (T419635)', diff saved to https://phabricator.wikimedia.org/P94117 and previous config saved to /var/cache/conftool/dbconfig/20260615-081440-fceratto.json
- 08:10 atsuko@deploy1003: Finished scap sync-world: Backport for translate: production opensearch on k8s endpoints (T425377) (duration: 20m 54s)
- 08:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 08:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Migration of es2047.codfw.wmnet completed
- 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 08:04 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1163', diff saved to https://phabricator.wikimedia.org/P94115 and previous config saved to /var/cache/conftool/dbconfig/20260615-080432-fceratto.json
- 08:03 atsuko@deploy1003: atsuko: Continuing with deployment
- 07:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 07:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 07:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 07:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 07:54 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1163', diff saved to https://phabricator.wikimedia.org/P94114 and previous config saved to /var/cache/conftool/dbconfig/20260615-075425-fceratto.json
- 07:53 atsuko@deploy1003: atsuko: Backport for translate: production opensearch on k8s endpoints (T425377) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 07:49 atsuko@deploy1003: Started scap sync-world: Backport for translate: production opensearch on k8s endpoints (T425377)
- 07:47 dcausse@deploy1003: mwscript-k8s job started: namespaceDupes cswiki --fix # T428619
- 07:46 dcausse@deploy1003: Finished scap sync-world: Backport for Switch wmgUseCalendar to false for dewikivoyage (T429095), Add alias namespace for cswiki (T428619) (duration: 34m 37s)
- 07:44 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1163 (T419635)', diff saved to https://phabricator.wikimedia.org/P94112 and previous config saved to /var/cache/conftool/dbconfig/20260615-074417-fceratto.json
- 07:43 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 07:39 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 07:33 dcausse@deploy1003: vadymts1, dcausse: Continuing with deployment
- 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 07:31 cwilliams@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:05:00 on db-test2001.codfw.wmnet with reason: Testing
- 07:28 dcausse@deploy1003: vadymts1, dcausse: Backport for Switch wmgUseCalendar to false for dewikivoyage (T429095), Add alias namespace for cswiki (T428619) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 07:26 elukey@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1080.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:26 elukey@cumin2002: START - Cookbook sre.hosts.provision for host cloudvirt1080.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:25 elukey@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:24 arnaudb@dns1005: END - running authdns-update
- 07:24 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db1163 (T419635)', diff saved to https://phabricator.wikimedia.org/P94110 and previous config saved to /var/cache/conftool/dbconfig/20260615-072446-fceratto.json
- 07:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1163.eqiad.wmnet with reason: Maintenance
- 07:24 elukey@cumin2002: START - Cookbook sre.hosts.provision for host cloudvirt1079.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:23 elukey@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:23 elukey@cumin2002: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Migration of es2047.codfw.wmnet completed
- 07:23 arnaudb@dns1005: START - running authdns-update
- 07:21 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 07:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:20 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 07:11 elukey@cumin2002: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
- 07:11 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2047.codfw.wmnet with OS trixie
- 07:11 dcausse@deploy1003: Started scap sync-world: Backport for Switch wmgUseCalendar to false for dewikivoyage (T429095), Add alias namespace for cswiki (T428619)
- 07:10 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 06:55 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2047.codfw.wmnet with reason: host reimage
- 06:53 moritzm: imported zookeeper 3.4.13-6+wmf12u1 to component/zookeeper34 for bookworm-wikimedia T428495
- 06:47 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2047.codfw.wmnet with reason: host reimage
- 06:31 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2047.codfw.wmnet with OS trixie
- 06:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2047: Upgrading es2047.codfw.wmnet
- 06:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2047: Upgrading es2047.codfw.wmnet
- 06:27 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 06:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2021: Migration to 10.11.18 T428861
- 06:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 06:09 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 06:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc2021: Migration to 10.11.18 T428861
- 05:59 marostegui: install mariadb 10.11.18 on pc1 T428861
- 05:57 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pc2021.codfw.wmnet,pc1021.eqiad.wmnet with reason: upgrading
- 05:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2021: Migration to 10.11.18 T428861
- 05:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 05:56 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 05:56 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2021: Migration to 10.11.18 T428861
- 05:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2021: Migration to 10.11.18 T428861
- 05:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc2021: Migration to 10.11.18 T428861
- 05:48 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Migration to 10.11.18 T428861
- 05:48 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Migration to 10.11.18 T428861
- 05:34 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es2046', diff saved to https://phabricator.wikimedia.org/P94105 and previous config saved to /var/cache/conftool/dbconfig/20260615-053403-marostegui.json
- 05:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on es2046.codfw.wmnet with reason: cloning
- 05:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on es2045.codfw.wmnet with reason: crash
- 05:30 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es2046', diff saved to https://phabricator.wikimedia.org/P94104 and previous config saved to /var/cache/conftool/dbconfig/20260615-053041-marostegui.json
- 02:18 Amir1: making Dexbot a bot in cywiki (T428927)
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s)
- 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-14
- 11:03 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 11:02 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 11:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 11:02 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 02:06 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-13
- 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s)
- 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2026-06-12
- 19:54 dwisehaupt@dns1004: END - running authdns-update
- 19:52 dwisehaupt@dns1004: START - running authdns-update
- 18:33 dwisehaupt@dns1006: END - running authdns-update
- 18:32 dwisehaupt@dns1006: START - running authdns-update
- 16:36 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 16:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 16:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 16:10 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 16:10 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 15:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 15:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 15:47 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 14:43 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for Hotfix for T428620 (T428620) (duration: 11m 17s)
- 14:36 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment
- 14:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for Hotfix for T428620 (T428620) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 14:31 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for Hotfix for T428620 (T428620)
- 14:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 14:28 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 13:24 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 13:24 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 12:22 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 12:22 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 12:22 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 12:22 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 12:17 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 12:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 12:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
- 12:04 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 12:04 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 12:04 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 12:03 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 12:02 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.changedisk (exit_code=99) for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 12:01 jmm@cumin2002: START - Cookbook sre.ganeti.changedisk for changing disk type of prometheus5003.eqsin.wmnet to drbd
- 11:40 moritzm: installing Linux 5.10.257 on Bullseye hosts
- 11:36 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 11:35 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 11:24 jelto@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Upgrade GitLab
- 11:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:56 atsuko@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
- 10:56 atsuko@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
- 10:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 10:49 atsuko@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
- 10:49 atsuko@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
- 10:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply
- 10:37 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply
- 10:36 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply
- 10:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply
- 10:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply
- 10:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply
- 10:12 atsuko@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
- 10:12 atsuko@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
- 10:08 jelto@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Upgrade GitLab
- 09:59 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
- 09:58 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
- 09:57 gkyziridis@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
- 06:13 jmm@cumin2002: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
- 06:11 jmm@cumin2002: START - Cookbook sre.puppet.disable-merges
- 03:07 ryankemper: T427951 sorry, `[eqiad,codfw].mediawiki.page_html_content_change.rc0` (accidentally a word)
- 03:06 ryankemper: T427951 Deleted all 20 unused dev/test topics on kafka-jumbo (verified empty first); 2 (`[eqiad,codfw]page_html_content_change.rc0`) were immediately auto-recreated empty by a still-running `dse-k8s` enrichment consumer; awaiting owner confirmation before final re-delete
- 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 13s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 00:00 bblack@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload and not P{cp7008.magru.wmnet} and A:cp - Upgrade wmfuniq to 0.3.0 ()
2026-06-11
- 22:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 22:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 22:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 22:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 22:05 egardner@deploy1003: Finished scap sync-world: Backport for Restore MediaViewer toggle in Special:Preferences (T428742) (duration: 30m 51s)
- 21:58 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host releases2003.codfw.wmnet with OS trixie
- 21:52 egardner@deploy1003: egardner: Continuing with deployment
- 21:51 egardner@deploy1003: egardner: Backport for Restore MediaViewer toggle in Special:Preferences (T428742) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:34 egardner@deploy1003: Started scap sync-world: Backport for Restore MediaViewer toggle in Special:Preferences (T428742)
- 21:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet with reason: host reimage
- 21:29 arlolra@deploy1003: Finished scap sync-world: Backport for Avoid the escaping from nowiki processing (T398967) (duration: 09m 09s)
- 21:28 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on releases2003.codfw.wmnet with reason: host reimage
- 21:25 arlolra@deploy1003: arlolra: Continuing with deployment
- 21:22 arlolra@deploy1003: arlolra: Backport for Avoid the escaping from nowiki processing (T398967) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 21:20 arlolra@deploy1003: Started scap sync-world: Backport for Avoid the escaping from nowiki processing (T398967)
- 21:07 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for badlogin for all small wikis (T426875), RadioRangeBallot: Fix strict mode issue (T428947) (duration: 10m 43s)
- 21:06 bblack@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text and not P{cp7008*} and A:cp - Upgrade wmfuniq to 0.3.0 ()
- 21:01 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 21:00 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for badlogin for all small wikis (T426875), RadioRangeBallot: Fix strict mode issue (T428947) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 20:56 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for badlogin for all small wikis (T426875), RadioRangeBallot: Fix strict mode issue (T428947)
- 20:51 jdrewniak@deploy1003: Finished scap sync-world: Backport for Donor Delight Badge: Unify on "Remove badge" language across treatments (T427313), [A11y] Donor Badge: Remove Badge button disappears too quickly (T428646), Donor Delight Badge, styles: Amending to final design review feedback (T427313) (duration: 34m 10s)
- 20:39 jdrewniak@deploy1003: annet, jdrewniak: Continuing with deployment
- 20:35 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host releases2003.codfw.wmnet with OS trixie
- 20:34 jdrewniak@deploy1003: annet, jdrewniak: Backport for Donor Delight Badge: Unify on "Remove badge" language across treatments (T427313), [A11y] Donor Badge: Remove Badge button disappears too quickly (T428646), Donor Delight Badge, styles: Amending to final design review feedback (T427313) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug
- 20:17 jdrewniak@deploy1003: Started scap sync-world: Backport for Donor Delight Badge: Unify on "Remove badge" language across treatments (T427313), [A11y] Donor Badge: Remove Badge button disappears too quickly (T428646), Donor Delight Badge, styles: Amending to final design review feedback (T427313)
- 19:12 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.6 refs T423915
- 18:12 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 18:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 17:52 reedy@deploy1003: Finished scap sync-world: Backport for UploadWizard.config.php: Fix cc-by-4.0-heirs msg issue (T428935 T405146) (duration: 08m 15s)
- 17:48 reedy@deploy1003: reedy: Continuing with deployment
- 17:46 reedy@deploy1003: reedy: Backport for UploadWizard.config.php: Fix cc-by-4.0-heirs msg issue (T428935 T405146) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 17:44 reedy@deploy1003: Started scap sync-world: Backport for UploadWizard.config.php: Fix cc-by-4.0-heirs msg issue (T428935 T405146)
- 17:26 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
- 17:25 blake@deploy1003: Scap cancelled without rolling back.
- 17:25 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
- 17:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
- 17:24 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
- 17:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
- 17:24 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
- 17:23 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
- 17:23 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
- 17:23 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
- 17:23 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
- 17:23 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
- 17:23 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
- 17:20 blake@deploy1003: blake: apache config update (T428772) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 17:20 blake@deploy1003: Started scap sync-world: apache config update (T428772)
- 17:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Migration of db2212.codfw.wmnet completed
- 17:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 17:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Migration of db1235.eqiad.wmnet completed
- 17:08 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
- 16:45 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:43 dzahn@dns1005: END - running authdns-update
- 16:42 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 16:41 dzahn@dns1005: START - running authdns-update
- 16:41 mutante: releases.wikimedia.org - switching backend from codfw to eqiad - releases1003 is now the source of rsync for uploaded releases files (use releases.discovery.wmnet to not have to think about it) - T418299
- 16:35 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts rdb2007.codfw.wmnet
- 16:35 jiji@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
- 16:35 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts rdb1011.eqiad.wmnet
- 16:35 jiji@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
- 16:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts rdb2009.codfw.wmnet
- 16:34 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 16:34 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: rdb2009.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jiji@cumin1003"
- 16:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Migration of db2212.codfw.wmnet completed
- 16:27 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: rdb2009.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jiji@cumin1003"
- 16:27 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Migration of db1235.eqiad.wmnet completed
- 16:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2212.codfw.wmnet with OS trixie
- 16:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1235.eqiad.wmnet with OS trixie
- 16:13 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 16:07 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 16:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 16:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 16:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 16:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 16:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2212.codfw.wmnet with reason: host reimage
- 16:01 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
- 16:01 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 16:01 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
- 16:01 kamila@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
- 16:00 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
- 16:00 kamila@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
- 16:00 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
- 16:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2212.codfw.wmnet with reason: host reimage
- 15:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1235.eqiad.wmnet with reason: host reimage
- 15:58 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 15:58 kamila@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
- 15:57 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
- 15:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 15:57 kamila@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
- 15:57 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
- 15:56 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb2009.codfw.wmnet
- 15:55 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 15:55 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb1011.eqiad.wmnet
- 15:55 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 15:55 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb2007.codfw.wmnet
- 15:54 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
- 15:54 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1235.eqiad.wmnet with reason: host reimage
- 15:54 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
- 15:53 kamila@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
- 15:53 kamila@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
- 15:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 15:40 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2212.codfw.wmnet with OS trixie
- 15:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 15:39 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1235.eqiad.wmnet with OS trixie
- 15:36 hnowlan@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 15:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1235: Upgrading db1235.eqiad.wmnet
- 15:35 hnowlan@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 15:35 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1235: Upgrading db1235.eqiad.wmnet
- 15:35 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 15:32 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:31 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
- 15:30 cscott@deploy1003: Finished scap sync-world: Backport for T428849: temporarily disable noisy warnings in HandleParsoidSectionLinks (T428849 T417530) (duration: 11m 29s)
- 15:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2212: Upgrading db2212.codfw.wmnet
- 15:26 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2212: Upgrading db2212.codfw.wmnet
- 15:26 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:26 cscott@deploy1003: cscott: Continuing with deployment
- 15:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1235: Upgrading db1235.eqiad.wmnet
- 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1235: Upgrading db1235.eqiad.wmnet
- 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 15:21 cscott@deploy1003: cscott: Backport for T428849: temporarily disable noisy warnings in HandleParsoidSectionLinks (T428849 T417530) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 15:19 cscott@deploy1003: Started scap sync-world: Backport for T428849: temporarily disable noisy warnings in HandleParsoidSectionLinks (T428849 T417530)
- 15:18 hnowlan@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 15:17 hnowlan@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 15:13 hnowlan@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 15:13 hnowlan@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 15:13 moritzm: installing libdbi-perl security updates
- 14:53 moritzm: installing Bind security updates (just client-side tools/libraries)
- 14:51 jmm@cumin2002: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-docker-registry (exit_code=0) rolling restart_daemons on A:docker-registry
- 14:48 jmm@cumin2002: START - Cookbook sre.misc-clusters.roll-restart-reboot-docker-registry rolling restart_daemons on A:docker-registry
- 14:43 moritzm: installing Poppler security updates
- 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 14:33 hnowlan@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 14:32 hnowlan@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 14:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 14:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Migration of db1234.eqiad.wmnet completed
- 14:26 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti5006.eqsin.wmnet to cluster eqsin02 and group 01
- 14:24 jmm@cumin2002: START - Cookbook sre.ganeti.addnode for new host ganeti5006.eqsin.wmnet to cluster eqsin02 and group 01
- 14:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 14:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply
- 14:18 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet
- 14:08 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet
- 14:00 Lucas_WMDE: UTC afternoon backport+config window done
- 13:58 javiermonton@deploy1003: Finished scap sync-world: Backport for stream: webrequest.page_view_stats.dev0 (T428725) (duration: 08m 12s)
- 13:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp5024.*
- 13:55 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5024.*
- 13:55 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp5020.*
- 13:54 javiermonton@deploy1003: javiermonton: Continuing with deployment
- 13:52 javiermonton@deploy1003: javiermonton: Backport for stream: webrequest.page_view_stats.dev0 (T428725) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:51 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs5004*} and A:liberica
- 13:50 javiermonton@deploy1003: Started scap sync-world: Backport for stream: webrequest.page_view_stats.dev0 (T428725)
- 13:50 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P{lvs5004*} and A:liberica
- 13:50 slyngs: reloading liberica config on lvs5004
- 13:50 moritzm: installing openssl security updates
- 13:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 13:46 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
- 13:46 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti5006.eqsin.wmnet with OS bookworm
- 13:46 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Migration of db1234.eqiad.wmnet completed
- 13:46 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
- 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
- 13:45 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
- 13:44 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2202.codfw.wmnet with OS trixie
- 13:43 alexsanford@deploy1003: Finished scap sync-world: Backport for Add 2FA enforcement demotion config for phase 3 groups (T423120) (duration: 07m 19s)
- 13:39 alexsanford@deploy1003: alexsanford: Continuing with deployment
- 13:38 alexsanford@deploy1003: alexsanford: Backport for Add 2FA enforcement demotion config for phase 3 groups (T423120) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:36 alexsanford@deploy1003: Started scap sync-world: Backport for Add 2FA enforcement demotion config for phase 3 groups (T423120)
- 13:36 slyngshede@dns1004: END - running authdns-update
- 13:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1234.eqiad.wmnet with OS trixie
- 13:34 moritzm: installing dovecot security updates
- 13:34 slyngshede@dns1004: START - running authdns-update
- 13:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 13:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for hCaptcha: Enable for MobileFrontend on all group1 wikis (T425940) (duration: 06m 59s)
- 13:29 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 13:29 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 13:29 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 13:29 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 13:28 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 13:28 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 13:28 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 13:27 dreamyjazz@deploy1003: dreamyjazz: Backport for hCaptcha: Enable for MobileFrontend on all group1 wikis (T425940) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: host reimage
- 13:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for hCaptcha: Enable for MobileFrontend on all group1 wikis (T425940)
- 13:25 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=mediawikiwiki '--reason=per phab:T428900' Wikimedia_Apps/Android_FAQ 'Wikimedia Apps/FAQ/Android' 'Martin Urbanec (WMF)' # T428900
- 13:24 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=mediawikiwiki '--reason=per phab:T428900' Wikimedia_Apps/Android_FAQ 'Wikimedia Apps/FAQ/Android' 'Martin Urbanec (WMF)' # T428900
- 13:22 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for fix: correct intake-url and payload type for NCS experiment events (T422295) (duration: 06m 51s)
- 13:22 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti5006.eqsin.wmnet with reason: host reimage
- 13:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1234.eqiad.wmnet with reason: host reimage
- 13:18 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, migr: Continuing with deployment
- 13:18 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2202.codfw.wmnet with reason: host reimage
- 13:18 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, migr: Backport for fix: correct intake-url and payload type for NCS experiment events (T422295) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:18 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 13:17 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 13:16 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for fix: correct intake-url and payload type for NCS experiment events (T422295)
- 13:15 jmm@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti5006.eqsin.wmnet with reason: host reimage
- 13:14 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=mediawikiwiki '--reason=per phab:T428900' Wikimedia_Apps/Android_FAQ 'Wikimedia Apps/FAQ/Android' 'Martin Urbanec (WMF)' # T428900
- 13:13 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 13:13 gkyziridis@deploy1003: Finished scap sync-world: Backport for wgRestSandboxSpecs: Add Lift Wing API to documentation wikis (T427902) (duration: 08m 47s)
- 13:13 andrewbogott: sudo -i reprepro --noskipold --component thirdparty/openstack-trixie-flamingo-backports update trixie-wikimedia
- 13:12 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1234.eqiad.wmnet with reason: host reimage
- 13:12 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 13:12 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=mediawikiwiki '--reason=per phab:T428900' Wikimedia_Apps/iOS_FAQ 'Wikimedia Apps/FAQ/iOS' 'Martin Urbanec (WMF)' # T428900
- 13:12 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 13:12 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 13:11 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 13:11 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 13:11 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 13:11 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
- 13:11 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
- 13:10 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
- 13:10 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
- 13:09 gkyziridis@deploy1003: gkyziridis: Continuing with deployment
- 13:06 gkyziridis@deploy1003: gkyziridis: Backport for wgRestSandboxSpecs: Add Lift Wing API to documentation wikis (T427902) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 13:06 claime: echo 'https://api.wikimedia.org/service/lw/specs/openapi.yaml' | mwscript-k8s --attach -- purgeList.php
- 13:04 gkyziridis@deploy1003: Started scap sync-world: Backport for wgRestSandboxSpecs: Add Lift Wing API to documentation wikis (T427902)
- 13:02 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2202.codfw.wmnet with OS trixie
- 13:00 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:57 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1234.eqiad.wmnet with OS trixie
- 12:55 moritzm: installing Exim security updates on Bullseye
- 12:47 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ganeti5006
- 12:47 jmm@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ganeti5006
- 12:46 jmm@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ganeti5006
- 12:46 jmm@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ganeti5006.eqsin.wmnet 9.0.132.10.in-addr.arpa 9.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 12:46 jmm@cumin2002: START - Cookbook sre.dns.wipe-cache ganeti5006.eqsin.wmnet 9.0.132.10.in-addr.arpa 9.0.0.0.0.0.0.0.2.3.1.0.0.1.0.0.1.0.1.0.0.0.5.e.2.f.d.0.1.0.0.2.ip6.arpa on all recursors
- 12:46 jmm@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 12:46 jmm@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ganeti5006 - jmm@cumin2002"
- 12:46 jmm@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ganeti5006 - jmm@cumin2002"
- 12:44 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1234: Upgrading db1234.eqiad.wmnet
- 12:44 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1234: Upgrading db1234.eqiad.wmnet
- 12:44 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 12:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Migration of db2188.codfw.wmnet completed
- 12:29 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "UX improvements - oblivian@cumin1003"
- 12:29 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: UX improvements - oblivian@cumin1003
- 12:28 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 12:28 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Migration of db1232.eqiad.wmnet completed
- 12:28 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: UX improvements - oblivian@cumin1003
- 12:28 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "UX improvements - oblivian@cumin1003"
- 12:27 jmm@cumin2002: START - Cookbook sre.dns.netbox
- 12:26 jmm@cumin2002: START - Cookbook sre.hosts.move-vlan for host ganeti5006
- 12:26 jmm@cumin2002: START - Cookbook sre.hosts.reimage for host ganeti5006.eqsin.wmnet with OS bookworm
- 12:21 moritzm: remove ganeti5006 from eqsin cluster for reimage T428229
- 12:17 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet
- 12:10 moritzm: installing openjdk-21 security updates on Bookworm
- 12:03 urbanecm@deploy1003: Finished scap sync-world: Backport for Remove GrowthExperiments extension from closed wikis (T428884) (duration: 06m 53s)
- 11:59 urbanecm@deploy1003: urbanecm: Continuing with deployment
- 11:58 urbanecm@deploy1003: urbanecm: Backport for Remove GrowthExperiments extension from closed wikis (T428884) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:56 urbanecm@deploy1003: Started scap sync-world: Backport for Remove GrowthExperiments extension from closed wikis (T428884)
- 11:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts rdb1012.eqiad.wmnet
- 11:49 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts rdb2010.codfw.wmnet
- 11:49 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:48 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: rdb2010.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jiji@cumin1003"
- 11:46 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts rdb2008.codfw.wmnet
- 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 11:46 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Migration of db2188.codfw.wmnet completed
- 11:44 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
- 11:43 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 11:43 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: rdb2010.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jiji@cumin1003"
- 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Migration of db1232.eqiad.wmnet completed
- 11:38 jiji@cumin1003: START - Cookbook sre.dns.netbox
- 11:37 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
- 11:37 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
- 11:36 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
- 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2188.codfw.wmnet with OS trixie
- 11:35 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb1012.eqiad.wmnet
- 11:34 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb2008.codfw.wmnet
- 11:34 jiji@cumin1003: START - Cookbook sre.hosts.decommission for hosts rdb2010.codfw.wmnet
- 11:33 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
- 11:32 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
- 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1232.eqiad.wmnet with OS trixie
- 11:27 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet
- 11:25 dreamyjazz@deploy1003: Finished scap sync-world: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on all wikis (T426039) (duration: 08m 38s)
- 11:21 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 11:19 dreamyjazz@deploy1003: dreamyjazz: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on all wikis (T426039) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 11:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2188.codfw.wmnet with reason: host reimage
- 11:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on all wikis (T426039)
- 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2188.codfw.wmnet with reason: host reimage
- 11:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1232.eqiad.wmnet with reason: host reimage
- 11:13 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet
- 11:13 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
- 11:12 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet
- 11:12 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet
- 11:11 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
- 11:09 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2001.codfw.wmnet
- 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1232.eqiad.wmnet with reason: host reimage
- 11:08 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet
- 11:05 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 11:04 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc2001.codfw.wmnet
- 11:04 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet
- 11:04 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 11:02 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1262.eqiad.wmnet with reason: crash
- 11:00 hnowlan@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
- 11:00 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet
- 10:59 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 10:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 10:58 hnowlan@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
- 10:55 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2188.codfw.wmnet with OS trixie
- 10:52 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: Upgrading db2188.codfw.wmnet
- 10:52 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2188: Upgrading db2188.codfw.wmnet
- 10:52 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:52 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1232.eqiad.wmnet with OS trixie
- 10:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1232: Upgrading db1232.eqiad.wmnet
- 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1232: Upgrading db1232.eqiad.wmnet
- 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 10:40 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:40 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 10:33 daniel@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
- 10:32 daniel@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
- 10:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on group 1 wikis (T426039) (duration: 11m 01s)
- 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
- 10:23 daniel@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
- 10:23 daniel@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
- 10:22 dreamyjazz@deploy1003: dreamyjazz: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on group 1 wikis (T426039) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 10:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for HCaptcha: Return 'forceshowcaptcha' error when CAPTCHA forced (T426476), hCaptcha: Enable for DiscussionTools on group 1 wikis (T426039)
- 10:18 daniel@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
- 10:18 daniel@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
- 10:10 hnowlan@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
- 10:10 hnowlan@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
- 10:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es2045.codfw.wmnet with OS trixie
- 10:09 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:08 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 10:06 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:05 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 10:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
- 10:02 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es2046', diff saved to https://phabricator.wikimedia.org/P94069 and previous config saved to /var/cache/conftool/dbconfig/20260611-100221-marostegui.json
- 10:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es2046', diff saved to https://phabricator.wikimedia.org/P94068 and previous config saved to /var/cache/conftool/dbconfig/20260611-100145-marostegui.json
- 10:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
- 09:59 jiji@deploy1003: Finished scap sync-world: Backport for ProductionServices.php: switch filebackend.php back to rdb1013 (T291916 T419976) (duration: 15m 41s)
- 09:54 jiji@deploy1003: jiji: Continuing with deployment
- 09:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es2045.codfw.wmnet with reason: host reimage
- 09:45 jiji@deploy1003: jiji: Backport for ProductionServices.php: switch filebackend.php back to rdb1013 (T291916 T419976) synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
- 09:43 jiji@deploy1003: Started scap sync-world: Backport for ProductionServices.php: switch filebackend.php back to rdb1013 (T291916 T419976)
- 09:42 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es2045.codfw.wmnet with reason: host reimage
- 09:37 elukey: uploaded spicerack_12.8.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
- 09:26 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2045.codfw.wmnet with OS trixie
- 09:26 marostegui@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host es2045.codfw.wmnet with OS bookworm
- 09:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 09:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Migration of db2176.codfw.wmnet completed
- 09:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
- 09:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Migration of db1219.eqiad.wmnet completed
- 09:11 claime: cumin -x 'A:swift-fe' "disable-puppet 'Disabling puppet for ratelimit deploy - cgoubert'"
- 08:57 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2045.codfw.wmnet with OS bookworm
- 08:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Migration of db2176.codfw.wmnet completed
- 08:34 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-test # T425377 populating ttmserver index on test cluster to estimate time required for the release (dblist: https://phabricator.wikimedia.org/P94055)
- 08:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Migration of db1219.eqiad.wmnet completed
- 08:33 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-test # T425377 populating ttmserver index on test cluster to estimate time required for the release (dblist: https://phabricator.wikimedia.org/P94053)
- 08:30 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@6200ab1] (releasing): T428823 (duration: 01m 18s)
- 08:29 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@6200ab1] (releasing): T428823
- 08:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2176.codfw.wmnet with OS trixie
- 08:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Migration to 10.11.17
- 08:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Migration to 10.11.17
- 08:25 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-test # T425377 populating ttmserver index on test cluster to estimate time required for the release (dblist: https://phabricator.wikimedia.org/P94052)
- 08:24 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@6200ab1] (releasing): Testing upgrade for T428823 (duration: 01m 17s)
- 08:23 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@6200ab1] (releasing): Testing upgrade for T428823
- 08:22 atsuko@deploy1003: mwscript-k8s job started: foreachwikiindblist mwscript.dblist extensions/Translate/scripts/ttmserver-export.php --ttmserver eqiad-test # T425377 populating ttmserver index on test cluster to estimate time required for the release (dblist: https://phabricator.wikimedia.org/P94051)
- 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1219.eqiad.wmnet with OS trixie
- 08:17 moritzm: installing PHP 8.2 security updates
- 08:15 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 08:14 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 08:11 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 08:11 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2176.codfw.wmnet with reason: host reimage
- 08:08 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb1013.eqiad.wmnet with OS trixie
- 08:06 jmm@cumin2002: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti5004.eqsin.wmnet to cluster eqsin02 and group 01
- 08:06 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 08:06 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 08:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pc2021.codfw.wmnet,pc1021.eqiad.wmnet with reason: upgrade
- 08:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1219.eqiad.wmnet with reason: host reimage
- 08:05 jmm@cumin2002: START - Cookbook sre.ganeti.addnode for new host ganeti5004.eqsin.wmnet to cluster eqsin02 and group 01
- 08:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Migration to 10.11.17 T427345
- 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Migration to 10.11.17 T427345
- 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2176.codfw.wmnet with reason: host reimage
- 08:04 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Migration to 10.11.17 T427345
- 08:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
- 08:03 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
- 08:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Migration to 10.11.17 T427345
- 07:59 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet
- 07:58 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1219.eqiad.wmnet with reason: host reimage
- 07:56 marostegui: install mariadb 10.11.17 on pc1 T427345
- 07:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb1013.eqiad.wmnet with reason: host reimage
- 07:50 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb1013.eqiad.wmnet with reason: host reimage
- 07:49 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 07:49 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 07:49 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet
- 07:47 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
- 07:47 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
- 07:46 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2176.codfw.wmnet with OS trixie
- 07:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1219.eqiad.wmnet with OS trixie
- 07:43 moritzm: imported Jenkins 2.541.3 for thirdparty/ci (Bullseye) and thirdparty/jenkins (Bookworm, Trixie)
- 07:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
- 07:35 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host rdb1013.eqiad.wmnet with OS trixie
- 07:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2176: Upgrading db2176.codfw.wmnet
- 07:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1219: Upgrading db1219.eqiad.wmnet
- 07:31 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2176: Upgrading db2176.codfw.wmnet
- 07:31 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 07:31 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1219: Upgrading db1219.eqiad.wmnet
- 07:31 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
- 07:31 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 07:30 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
- 07:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Repooling
- 07:19 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
- 06:51 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2045.codfw.wmnet with OS trixie
- 06:50 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es2042', diff saved to https://phabricator.wikimedia.org/P94044 and previous config saved to /var/cache/conftool/dbconfig/20260611-065049-marostegui.json
- 06:50 marostegui@cumin1003: dbctl commit (dc=all): 'Depool es2042', diff saved to https://phabricator.wikimedia.org/P94043 and previous config saved to /var/cache/conftool/dbconfig/20260611-065027-marostegui.json
- 06:44 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Repooling
- 06:43 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db1163 T426083', diff saved to https://phabricator.wikimedia.org/P94041 and previous config saved to /var/cache/conftool/dbconfig/20260611-064319-fceratto.json
- 06:42 fceratto@dns1005: END - running authdns-update
- 06:40 fceratto@dns1005: START - running authdns-update
- 06:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.global-read-only (exit_code=0)
- 06:33 fceratto@cumin1003: MariaDB change: Setting sections s1 as read-write for T426083: 'Maintenance until 06:15 UTC'
- 06:33 fceratto@cumin1003: START - Cookbook sre.mysql.global-read-only
- 06:33 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db1184 to s1 primary and set section read-write T426083', diff saved to https://phabricator.wikimedia.org/P94040 and previous config saved to /var/cache/conftool/dbconfig/20260611-063323-fceratto.json
- 06:32 fceratto@cumin1003: dbctl commit (dc=all): 'Set s1 eqiad as read-only for maintenance - T426083', diff saved to https://phabricator.wikimedia.org/P94039 and previous config saved to /var/cache/conftool/dbconfig/20260611-063251-fceratto.json
- 06:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.global-read-only (exit_code=0)
- 06:32 fceratto@cumin1003: Dbctl change: Setting sections s1 as read-write for T426083: 'Maintenance until 06:15 UTC'
- 06:32 fceratto@cumin1003: MariaDB change: Setting sections s1 as read-write for T426083: 'Maintenance until 06:15 UTC'
- 06:31 fceratto@cumin1003: START - Cookbook sre.mysql.global-read-only
- 06:31 fceratto@cumin1003: dbctl commit (dc=all): 'Set s1 eqiad as read-only for maintenance - T426083', diff saved to https://phabricator.wikimedia.org/P94037 and previous config saved to /var/cache/conftool/dbconfig/20260611-063100-fceratto.json
- 06:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.global-read-only (exit_code=0)
- 06:30 fceratto@cumin1003: MariaDB change: Setting sections s1 as read-only for T426083: 'Maintenance until 06:15 UTC'
- 06:30 fceratto@cumin1003: Dbctl change: Setting sections s1 as read-only for T426083: 'Maintenance until 06:15 UTC'
- 06:29 fceratto@cumin1003: START - Cookbook sre.mysql.global-read-only
- 06:29 federico3: Starting s1 eqiad failover from db1163 to db1184 - T426083
- 06:22 fceratto@cumin1003: dbctl commit (dc=all): 'Set db1184 with weight 0 T426083', diff saved to https://phabricator.wikimedia.org/P94035 and previous config saved to /var/cache/conftool/dbconfig/20260611-062224-fceratto.json
- 06:22 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s1 T426083
- 05:37 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
- 05:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
- 05:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
- 05:18 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
- 05:17 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es2045.codfw.wmnet with OS trixie
- 05:17 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2045: Upgrading es2045.codfw.wmnet
- 05:16 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2045: Upgrading es2045.codfw.wmnet
- 05:16 marostegui@cumin1003: START - Cookbook sre.mysql.major-upgrade
- 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 44s)
- 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
- 01:23 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.*
- 01:19 jasmine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
- 01:18 jasmine@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
- 01:18 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kafka-main1009.eqiad.wmnet with OS trixie
- 01:12 jasmine@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 01:12 jasmine@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
- 01:12 jasmine@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 01:12 jasmine@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
- 01:11 jasmine@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 01:11 jasmine@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 01:11 jasmine@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 01:10 jasmine@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 01:10 jasmine@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
- 01:09 jasmine@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
- 01:09 jasmine@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
- 01:08 jasmine@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
- 01:08 jasmine@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
- 01:08 jasmine@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
- 01:07 jasmine@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 01:07 jasmine@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 01:06 jasmine@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 01:06 jasmine@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 01:06 jasmine@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 01:05 jasmine@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 01:05 jasmine@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
- 01:05 jasmine@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
- 01:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kafka-main1009.eqiad.wmnet with reason: host reimage
- 00:58 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on kafka-main1009.eqiad.wmnet with reason: host reimage
- 00:54 jasmine@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 00:53 jasmine@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
- 00:53 jasmine@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 00:53 jasmine@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
- 00:53 jasmine@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
- 00:53 jasmine@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
- 00:52 jasmine@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
- 00:51 jasmine@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
- 00:41 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host kafka-main1009
- 00:41 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host kafka-main1009
- 00:41 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host kafka-main1009
- 00:41 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kafka-main1009.eqiad.wmnet 37.48.64.10.in-addr.arpa 7.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 00:41 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache kafka-main1009.eqiad.wmnet 37.48.64.10.in-addr.arpa 7.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
- 00:41 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
- 00:41 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host kafka-main1009 - jasmine@cumin2002"
- 00:40 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host kafka-main1009 - jasmine@cumin2002"
- 00:39 cdanis@cumin1003: dbctl commit (dc=all): 'depool db1262', diff saved to https://phabricator.wikimedia.org/P94032 and previous config saved to /var/cache/conftool/dbconfig/20260611-003950-cdanis.json
- 00:36 jasmine@cumin2002: START - Cookbook sre.dns.netbox
- 00:34 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp5020.*
- 00:30 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host kafka-main1009
- 00:30 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host kafka-main1009.eqiad.wmnet with OS trixie
- 00:03 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp5024.*
Other archives
2000s
- Archive 1: 2004 Jun - 2004 Sep
- Archive 2: 2004 Oct - 2004 Nov
- Archive 3: 2004 Dec - 2005 Mar
- Archive 4: 2005 Apr - 2005 Jul
- Archive 5: 2005 Aug - 2005 Oct, with revision history 2004-06-23 to 2005-11-25
- Archive 6: 2005 Nov - 2006 Feb
- Archive 7: 2006 Mar - 2006 Jun
- Archive 8: 2006 Jul - 2006 Sep
- Archive 9: 2006 Oct - 2007 Jan, with revision history 2005-11-25 to 2007-02-21
- Archive 10: 2007 Feb - 2007 Jun
- Archive 11: 2007 Jul - 2007 Dec
- Archive 12: 2008 Jan - 2008 Jul
- Archive 12a: 2008 Aug
- Archive 12b: 2008 Sept
- Archive 13: 2008 Oct - 2009 Jun
- Archive 14: 2009 Jun - 2009 Dec
2010s
- Archive 15: 2010 Jan - 2010 Jun
- Archive 16: 2010 Jul - 2010 Oct
- Archive 17: 2010 Nov - 2010 Dec
- Archive 18: 2011 Jan - 2011 Jun
- Archive 19: 2011 Jul - 2011 Dec
- Archive 20: 2011 Dec - 2012 Jun, with revision history 2007-02-21 to 2012-03-27
- Archive 21: 2012 Jul - 2013 Jan
- Archive 22: 2013 Jan - 2013 Jul
- Archive 23: 2013 Aug - 2013 Dec
- Archive 24: 2014 Jan - 2014 Mar
- Archive 25: 2014 April - 2014 September
- Archive 26: 2014 October - 2014 December
- Archive 27: 2015 January - 2015 July
- Archive 28: 2015 August - 2015 December
- Archive 29: 2016 January - 2016 May
- Archive 30: 2016 June - 2016 August
- Archive 31: 2016 September - 2016 December
- Archive 32: 2017 January - 2017 July
- Archive 33: 2017 August - 2017 December
- Archive 34: 2018 January - 2018 April
- Archive 35: 2018 May - 2018 August
- Archive 36: 2018 September - 2018 December
- Archive 37: 2019 January - 2019 April
- Archive 38: 2019 May - 2019 August
- Archive 39: 2019 September - 2019 December
2020-2024
- Archive 40: 2020 January - 2020 April
- Archive 41: 2020 May - 2020 July
- Archive 42: 2020 August - 2020 November
- Archive 43: 2020 December
- Archive 44: 2021 January - 2021 April
- Archive 45: 2021 May - 2021 July
- Archive 46: 2021 August - 2021 October
- Archive 47: 2021 November - 2021 December
- Archive 48: 2022 January
- Archive 49: 2022 February
- Archive 50: 2022 March
- Archive 51: 2022 April 1-15
- Archive 52: 2022 April 16-30
- Archive 53: 2022 May
- Archive 54: 2022 June
- Archive 55: 2022 July
- Archive 56: 2022 August
- Archive 57: 2022 September
- Archive 58: 2022 October
- Archive 59: 2022 November 1-15
- Archive 60: 2022 November 16-30
- Archive 61: 2022 December
- Archive 62: 2023 January
- Archive 63: 2023 February
- Archive 64: 2023 March
- Archive 65: 2023 April
- Archive 66: 2023 May
- Archive 67: 2023 June
- Archive 68: 2023 July
- Archive 69: 2023 August 1-15
- Archive 70: 2023 August 16-31
- Archive 71: 2023 September
- Archive 72: 2023 October
- Archive 73: 2023 November
- Archive 74: 2023 December
- Archive 75: 2024 January
- Archive 76: 2024 February
- Archive 77: 2024 March
- Archive 78: 2024 April
- Archive 79: 2024 May 1-15
- Archive 80: 2024 May 16-31
- Archive 81: 2024 June 1-15
- Archive 82: 2024 June 16-30
- Archive 83: 2024 July
- Archive 84: 2024 August
- Archive 85: 2024 September
- Archive 86: 2024 October
- Archive 87: 2024 November
- Archive 88: 2024 December
2025-present
- Archive 89: 2025 January
- Archive 90: 2025 February
- Archive 91: 2025 March
- Archive 92: 2025 April
- Archive 93: 2025 May
- Archive 94: 2025 June
- Archive 95: 2025 July
- Archive 96: 2025 August
- Archive 97: 2025 September
- Archive 98: 2025 October
- Archive 99: 2025 November
- Archive 100: 2025 December
- Archive 101: 2026 January
- Archive 102: 2026 February
- Archive 103: 2026 March
- Archive 104: 2026 April
- Archive 105: 2026 May 1-20
- Archive 106: 2026 May 21-June 10
- Archive 107: 2026 June 11-30