ATLAS Media Group Status Page

Some systems are experiencing issues.

ATLAS Media Group Status Page

The latest status updates from the team. Updates provided on known and ongoing incidents.

Past incidents

No incidents reported.

No incidents reported.

Fixed

6 months ago

Full service has been restored.

Reported

6 months ago

We're seeing intermittent outages on MastodonAppUK and are investigating the issue.

6 months ago
Fixed

Fixed

6 months ago

Full service has been restored.

Watching

6 months ago

Our backups are now in a better state, we're re-starting the Universeodon service at this time.

Identified

6 months ago

We're continuing to see a large backlog in WAL Files needing to be pushed to our external backup service. We're going to continue to monitor and we will look to keep the site offline until we've made a significant enough dent in the backlog to safely restore service and not risk backup stability.

Identified

6 months ago

Due to growing disk usage and our backups struggling to keep up right now we've temporarily taken Universeodon.com offline to allow the backups a chance to fully catch-up. Unfortunately once we increase the disk allocation to the DB Server it's impossible to reverse and with the current constraints with SSD Storage globally we are looking to conserve capacity where possible.

Identified

6 months ago

We can see a large backlog in WAL Files waiting to be pushed to our off-site backup service, this is likely the result of the issues with our networking and routing which has taken DNS offline a few times. We expect that once all these files are pushed up we should see a substantial amount of disk space released to the DB Server.

Investigating

6 months ago

Despite a significant increase in disk capacity we're once again seeing disk related issues impacting Universeodon - We suspect this might be the result of issues with our backup streaming. We are investigating.

Watching

6 months ago

We have expanded the disk space and will be reviewing our monitoring tool configuration to ensure high disk capacity warnings are flagged in future to allow us to proactively intervene. We will continue to monitor to ensure full stability has restored. As part of this we are also validating our off-site backup configuration has not been impacted and is still operational.

Identified

6 months ago

We have identified the issue as being a full disk on our primary database server. We're working to expand capacity now and restore service as quickly as possible.

Reported

6 months ago

Monitoring has detected a global outage of Universeodon - We are investigating.

No incidents reported.

Fixed

Fixed

6 months ago

Full service is restored and stable at this time.

Watching

6 months ago

We have remotely restarted the app server which ran into issues and it appears that has restored the connectivity that had been broken. We are still seeing issues with our gateway router and will continue updates on the other outstanding incident.

Identified

6 months ago

Our gateway has now failed in such a way that means we can't properly access the devices and will need to attempt to restore it's configuration from a recent backup. We're working to restore functionality as quickly as possible.

Investigating

6 months ago

We think at this time the issue may be related to the ongoing major issus with our router device. We're restarting the router now as it appears to have gone non responsive and we hope that will restore connectivity. We will monitor the situation at this time for the completion of the reboot and take further steps as appropriate.

Reported

6 months ago

Monitoring has detected one of our servers has gone offline. Support are investigating.

Fixed

6 months ago

We have provided additional diagnostics data to the vendor and continue to wait for their response. We will continue to look at mitigations and ultimately moving away from this vendors hardware however as we've been able to mitigate the active impact we're going to close this incident.

Identified

6 months ago

We are continuing to see random outages of our gateway device which have now resulted in further significant incidents and impact to service. We're continuing to work to resolve the situation.

Identified

7 months ago

It looks like the disruption we saw is related to the ongoing issues with our gateway router which caused a temporary interuption to traffic routing to and from our servers. We're attempting now to get access to the remote gateway to perform a restart which should resolve the situation.

Reported

7 months ago

We're seeing major networking disruption at our Redditch DC - We are investigating.