Thread

DW
Daniel Wroblewski9:03 AMOpen in Slack
Hi everyone, it's me again 😅 Am I the only one having performance issues with 1.4 RC's UI? We updated accidentally, because Dockerhub's "latest" tag doesnt care about the release being stable or not 🫠 Since then I kept updating it to the most recent RCs (currently rc31), but still the issues persist.
Problems:
1. The spinner in the sidebar keeps rotating for a 2-3 minutes before it stops (screenshot 1). THis is not a problem on its own, but I assume this may be related to some other issues
2. When I click Studio > Agents, nothing happens. Like literally, there is no feedback to the user that it has been clicked. And again after sometimes 30 seconds, sometimes 2 minutes, the Agents page finally load. Or sometimes they dont and nothing ever happens or I see errors (samples attached in screenshot 2 and 3)
3. MCP registry - same thing as with Agents, sometimes it loads, sometimes it doesnt. If it does, it takes a long time.
4. The entire web app feels laggy 😕 Hard to tell exactly, but the UX is just of lower quality than it was on 1.3, especially the menu items.
None of those issues existed on 1.3. I wish we could downgrade to 1.3, but I think there were some DB migrations in between and Archestra won't start after pinning the docker version to 1.3.x. Unless there is some easy way to rollback and maintain all the work we did on 1.4 (creating new agents).
Of course all caches have been cleared, and Archestra is used in a fresh browser icognito session. It sits behind Cloudflare, which also had its Cache purged many times. We only have 9 agents (including users' personal My Assistants) and 10 MCPs (9 self hosted, 1 external). CPU on this machine is at 50% most of the time, RAM at 60%
I am happy to provide whatever additional details might be needed.

13 replies
A
arseny9:06 AMOpen in Slack
hey @user, thanks for the report! 1.4 is pretty unstable at this point, we’ll take care of those defects asap and will report back. And yes, the latest tag is indeed a problem
A
arseny9:18 AMOpen in Slack
@user we’re reproducing this on our side too, but 3 quick things would help us narrow it down a lot:
1. docker exec <container> env | grep -E 'AGENT_RUNTIME|BETA' – are any of these set?
2. docker stats + docker exec <container> ps aux --sort=-%cpu | head -15 – which process is eating the CPU?
3. Browser DevTools → Network tab, while the sidebar spinner is turning: which request(s) are stuck in “pending”?
DW
Daniel Wroblewski9:40 AMOpen in Slack
@user THanks! Here are the outputs:
ARCHESTRA_AGENT_RUNTIME_ENABLED=true

root@idh-admin-archestra:~# docker exec archestra-mcp-control-plane env | grep -E 'AGENT_RUNTIME|BETA'
root@idh-admin-archestra:~#```
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS
e9189b871861 archestra 10.70% 1.356GiB / 6.807GiB 19.92% 503MB / 69.4MB 824MB / 800MB 65
661d311686b2 archestra-mcp-control-plane 65.35% 2.438GiB / 6.807GiB 35.82% 3.72GB / 3.77GB 4.41GB / 30.8GB 669
e7827acd02ad archestra-postgres 0.44% 411.7MiB / 6.807GiB 5.91% 17.9MB / 92.5MB 211MB / 74.6MB 17
3002c2f31c7f postgres-exporter 0.00% 17.52MiB / 6.807GiB 0.25% 0B / 0B 26.5MB / 1.93MB 8```
ps: unrecognized option: sort=-%cpu
BusyBox v1.37.0 (2025-12-16 14:19:28 UTC) multi-call binary.

Usage: ps [-o COL1,COL2=HEADER] [-T]

Show list of processes

        -o COL1,COL2=HEADER     Select columns for display
        -T                      Show threads

root@idh-admin-archestra:~# docker exec archestra ps aux  | head -15
PID   USER     TIME  COMMAND
    1 root      0:00 {docker-entrypoi} /bin/sh /docker-entrypoint.sh
  271 root      0:10 {supervisord} /usr/bin/python3 /usr/bin/supervisord -c /etc/supervisord.conf
  279 root      4:33 {MainThread} node --enable-source-maps --report-on-fatalerror --report-uncaught-exception --diagnostic-dir=/app/data/diagnostics dist/server.mjs
  280 root      1:37 next-server (v
  432 root      1:59 /usr/local/bin/dagger session --workdir / --label <http://dagger.io/sdk.name:rust|dagger.io/sdk.name:rust> --label <http://dagger.io/sdk.version:0.21.9|dagger.io/sdk.version:0.21.9>
  448 root      0:07 kubectl --context= --namespace=default exec --container=dagger-engine -i dagger-runtime-engine-0 -- buildctl dial-stdio
  457 root      0:04 kubectl --context= --namespace=default exec --container=dagger-engine -i dagger-runtime-engine-0 -- buildctl dial-stdio
  463 root      0:03 kubectl --context= --namespace=default exec --container=dagger-engine -i dagger-runtime-engine-0 -- buildctl dial-stdio
 1320 root      0:00 ps aux

root@idh-admin-archestra:~# docker exec archestra-mcp-control-plane ps aux --sort=-%cpu | head -15
USER         PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root     1258104  116  0.0   8096  4224 ?        Rs   09:35   0:00 ps aux --sort=-%cpu
root     1258090 55.8  0.7 1324532 55680 ?       Ssl  09:35   0:00 dagger core version
root         136 46.9  1.3 3200796 93216 ?       Ssl  Sep30 3420:42 /usr/local/bin/containerd
root        3753 14.7  3.2 1668352 229136 ?      Sl   Sep30 1076:38 /usr/local/bin/dagger-engine --config /etc/dagger/engine.toml
root         225 11.8  1.2 2472712 91032 ?       Ssl  Sep30 867:12 /usr/bin/kubelet --bootstrap-kubeconfig=/etc/kubernetes/bootstrap-kubelet.conf --kubeconfig=/etc/kubernetes/kubelet.conf --config=/var/lib/kubelet/config.yaml --node-ip=172.18.0.3 --node-labels= --pod-infra-container-image=<http://registry.k8s.io/pause:3.10.1|registry.k8s.io/pause:3.10.1> --provider-id=<kind://docker/archestra-mcp/archestra-mcp-control-plane> --runtime-cgroups=/system.slice/containerd.service
root         722 11.6  4.0 1518316 287796 ?      Ssl  Sep30 853:13 kube-apiserver --advertise-address=172.18.0.3 --allow-privileged=true --authorization-mode=Node,RBAC --client-ca-file=/etc/kubernetes/pki/ca.crt --enable-admission-plugins=NodeRestriction --enable-bootstrap-token-auth=true --etcd-cafile=/etc/kubernetes/pki/etcd/ca.crt --etcd-certfile=/etc/kubernetes/pki/apiserver-etcd-client.crt --etcd-keyfile=/etc/kubernetes/pki/apiserver-etcd-client.key --etcd-servers=<https://127.0.0.1:2379> --kubelet-client-certificate=/etc/kubernetes/pki/apiserver-kubelet-client.crt --kubelet-client-key=/etc/kubernetes/pki/apiserver-kubelet-client.key --kubelet-preferred-address-types=InternalIP,ExternalIP,Hostname --proxy-client-cert-file=/etc/kubernetes/pki/front-proxy-client.crt --proxy-client-key-file=/etc/kubernetes/pki/front-proxy-client.key --requestheader-allowed-names=front-proxy-client --requestheader-client-ca-file=/etc/kubernetes/pki/front-proxy-ca.crt --requestheader-extra-headers-prefix=X-Remote-Extra- --requestheader-group-headers=X-Remote-Group --requestheader-username-headers=X-Remote-User --runtime-config= --secure-port=6443 --service-account-issuer=<https://kubernetes.default.svc.cluster.local> --service-account-key-file=/etc/kubernetes/pki/sa.pub --service-account-signing-key-file=/etc/kubernetes/pki/sa.key --service-cluster-ip-range=10.96.0.0/16 --tls-cert-file=/etc/kubernetes/pki/apiserver.crt --tls-private-key-file=/etc/kubernetes/pki/apiserver.key
root         757  6.4  0.8 11740748 57716 ?      Ssl  Sep30 471:28 etcd --advertise-client-urls=<https://172.18.0.3:2379> --cert-file=/etc/kubernetes/pki/etcd/server.crt --client-cert-auth=true --data-dir=/var/lib/etcd --feature-gates=InitialCorruptCheck=true --initial-advertise-peer-urls=<https://172.18.0.3:2380> --initial-cluster=archestra-mcp-control-plane=<https://172.18.0.3:2380> --key-file=/etc/kubernetes/pki/etcd/server.key --listen-client-urls=<https://127.0.0.1:2379>,<https://172.18.0.3:2379> --listen-metrics-urls=<http://127.0.0.1:2381> --listen-peer-urls=<https://172.18.0.3:2380> --name=archestra-mcp-control-plane --peer-cert-file=/etc/kubernetes/pki/etcd/peer.crt --peer-client-cert-auth=true --peer-key-file=/etc/kubernetes/pki/etcd/peer.key --peer-trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt --snapshot-count=10000 --trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt --watch-progress-notify-interval=5s
root     1235379  5.3  1.3 1298760 98748 ?       Ssl  08:34   3:14 kube-controller-manager --allocate-node-cidrs=true --authentication-kubeconfig=/etc/kubernetes/controller-manager.conf --authorization-kubeconfig=/etc/kubernetes/controller-manager.conf --bind-address=127.0.0.1 --client-ca-file=/etc/kubernetes/pki/ca.crt --cluster-cidr=10.244.0.0/16 --cluster-name=archestra-mcp --cluster-signing-cert-file=/etc/kubernetes/pki/ca.crt --cluster-signing-key-file=/etc/kubernetes/pki/ca.key --controllers=*,bootstrapsigner,tokencleaner --enable-hostpath-provisioner=true --kubeconfig=/etc/kubernetes/controller-manager.conf --leader-elect=true --requestheader-client-ca-file=/etc/kubernetes/pki/front-proxy-ca.crt --root-ca-file=/etc/kubernetes/pki/ca.crt --service-account-private-key-file=/etc/kubernetes/pki/sa.key --service-cluster-ip-range=10.96.0.0/16 --use-service-account-credentials=true
root     1235470  2.6  0.6 1275956 47488 ?       Ssl  08:34   1:38 kube-scheduler --authentication-kubeconfig=/etc/kubernetes/scheduler.conf --authorization-kubeconfig=/etc/kubernetes/scheduler.conf --bind-address=127.0.0.1 --kubeconfig=/etc/kubernetes/scheduler.conf --leader-elect=true
65532       3541  1.0  0.3 1298924 23660 ?       Ssl  Sep30  74:11 /coredns -conf /etc/coredns/Corefile
1000       13077  0.9  1.7 157684 122964 ?       Rl   Sep30  72:05 /home/mcp/.cache/uv/archive-v0/q9t1dVRHiD0lczQ7yf1QY/bin/python /home/mcp/.cache/uv/archive-v0/q9t1dVRHiD0lczQ7yf1QY/bin/workspace-mcp --transport streamable-http --tools drive sheets
65532       3500  0.9  0.3 1299180 23544 ?       Ssl  Sep30  71:01 /coredns -conf /etc/coredns/Corefile
1000       12906  0.9  1.5 144296 107088 ?       Sl   Sep30  68:44 /home/mcp/.cache/uv/archive-v0/X7aUvdW8nqW_jJDQcF1yR/bin/python /home/mcp/.cache/uv/archive-v0/X7aUvdW8nqW_jJDQcF1yR/bin/workspace-mcp --transport streamable-http --tools docs
1000       13780  0.7  1.7 20077312 126196 ?     Sl   Sep30  56:38 node /home/mcp/.npm/_npx/1170ee3efa7f238b/node_modules/.bin/mcp-gitlab```
Those were the ones pending (attachment). Usually they all resolve around the same time so I cant tell the exact one
A
arseny10:24 AMOpen in Slack
Found one suspect: with ARCHESTRA_AGENT_RUNTIME_ENABLED=true, 1.4 constantly pre-pulls large agent images into the embedded k8s cluster, which hammers disk/CPU and can knock over the cluster’s control plane.
could you restart with ARCHESTRA_AGENT_RUNTIME_ENABLED=false (if you’re not actively using that feature yet) and tell me whether the UI slowness goes away too? If it doesn’t, please send:
docker logs archestra --since 15m 2>&1 | grep -iE 'responseTime|timeout|ECONNREFUSED' | tail -300```
DW
Daniel Wroblewski11:33 AMOpen in Slack
@user Thank you so much for your support with this so far. Disabling this feature flag didnt help (no changes).
You can see its disabled now:
ARCHESTRA_AGENT_RUNTIME_ENABLED=false```

Please see the logs attached. THere are some database timeouts, so maybe they are the main issue here
A
arseny11:41 AMOpen in Slack
@user Thanks, that’s exactly what we needed: the slow part is the new permission checks in 1.4 timing out in the DB. Could you run this so we can match your data shape?
select resource, count(*) as policies, sum(jsonb_array_length(grants)) as grants, max(jsonb_array_length(grants)) as max_grants
  from resource_permission_policies group by 1 order by 3 desc;
select (select count(*) from agents) as agents_total,
       (select count(*) from agents where deleted_at is null) as agents_live,
       (select count(*) from internal_mcp_catalog) as catalog_items,
       (select count(*) from team) as teams,
       (select count(*) from team where parent_team_id is not null) as nested_teams,
       (select count(*) from team_member) as team_members,
       (select count(*) from member) as members,
       (select count(*) from organization_role) as custom_roles;
select role from member where user_id = '1aj1WaJSFtb0gQE8BD343W1408PPDcl5';```
DW
Daniel Wroblewski11:47 AMOpen in Slack
Sure thing, here it is @user :
--------------------+----------+--------+------------
 llmModel           |      846 |   4168 |          5
 conversation       |      156 |    157 |          2
 agent              |       35 |     69 |          5
 mcpRegistry        |       27 |     62 |          4
 skill              |        9 |     22 |          4
 app                |        5 |     18 |          4
 llmProviderApiKey  |       10 |     12 |          2
 knowledgeConnector |        4 |     11 |          4
 knowledgeBase      |        3 |     10 |          4
 mcpGateway         |       10 |      9 |          2
 llmVirtualKey      |        4 |      5 |          2
 environment        |        1 |      3 |          3
 project            |        2 |      3 |          2
 plugin             |        1 |      2 |          2
 serviceAccount     |        1 |      2 |          2
 llmOauthClient     |        1 |      2 |          2
 mcpOauthClient     |        1 |      2 |          2
 scheduledTask      |        1 |      2 |          2
 knowledgeFile      |        1 |      2 |          2
 agentRun           |        1 |      1 |          1
(20 rows)

 agents_total | agents_live | catalog_items | teams | nested_teams | team_members | members | custom_roles
--------------+-------------+---------------+-------+--------------+--------------+---------+--------------
           44 |          30 |            32 |     5 |            0 |           11 |       7 |            0
(1 row)

 role
-------
 admin
(1 row)```
(takes less than 1 second btw)
A
arseny12:34 PMOpen in Slack
@user I think we found it: 1.4's new permission queries make Postgres JIT-compile them on every request (3-4s each, while the real work is 2ms), and under a page load’s ~40 parallel requests that snowballs into the 30s timeouts.
Could you check and try the workaround?
1. Check JIT is on:
2. If both say on/true, disable it (safe and reversible: `... SET jit = on` undoes it):
docker restart archestra```
Then tell us whether Agents / MCP Registry are fast again. The proper fix will come in the next RC.
DW
Daniel Wroblewski12:42 PMOpen in Slack
@user Yes, it was true/on. I altered it to off. Agents and MCP registries are now loading fast again, awesome! Thank you!
The postgres I am using is pgvector/pgvector:0.8.1-pg18-trixie
👀1
A
arseny12:44 PMOpen in Slack
nice! thanks for helping with the triage, i’ll bake a fix asap!
❤️1
DW
Daniel Wroblewski12:49 PMOpen in Slack
Love your responsiveness and helpfulness. As I mentioned earlier, we love the product too ❤️
@user Are those slack messages a good place for me to report stuff? Or would you guys rather have me submit a standard GitHub issue? Or maybe have GitHub issues for a small issues and report big blockers here on Slack? Whatever works for you best
A
arseny12:52 PMOpen in Slack
Slack is typically faster especially for non-trivial asks that may need iterations, so don’t hesitate to ping me or anyone from the team here
❤️1
A
arseny3:10 PMOpen in Slack
github.com/archestra-ai/archestra/pull/8434 is queued for long-term fix. And the one leftover is to separate latest and latest-stable.
❤️1