ATLAS Media Group Status Page

All systems are operational.

ATLAS Media Group Status Page

The latest status updates from the team. Updates provided on known and ongoing incidents.

Planned maintenance

  • We need to perform routine maintenance to our physical hardware located within the Redditch Data center. As part of this we will need to take all services offline for a short amount of time while we apply system and security updates to our shared storage pool and restart the underlying host. We will also be performing patching and upgrades to our application hosts during this time and will be failing over our applications across the nodes in the environment.

    We expect the maintenance itself to take approx 3 hours with a short outage of up to 30 mins throughout this while we take the environment offline for full system restarts and hardware improvements.

    Affected Components: VPS Control Panel, Universeodon Database, Universeodon Queue & Content Processing Service, Universeodon Website, Universeodon Relay, Billing Panel, Blog, MastodonApp.UK Website, MastodonAppUK Database Service and MastodonAppUK Queue & Content Processing Service

Past incidents

Fixed

1 year ago

Service restored.

Reported

1 year ago

We have identified an issue whereby our Nexus server is failing to automatically renew certificates as expected, the site is currently offline as a result.

Fixed

1 year ago

Full service has now been restored.

Identified

1 year ago

Storage has been restarted. Some services have come online however we are still trying to get all services online.

Reported

1 year ago

We are investigating an issue impacting multiple services due to shared storage becoming unavailable. We will update as soon as we know more.

Fixed

1 year ago

We are back fully online.

Identified

1 year ago

We have identified the issue on both the old and new host and the original host is now fully operational. There was an additional issue when we attempted to migrate clients onto a new host, we are currently reversing the migrations now.

Reported

1 year ago

We are working to resolve issues with one of the nodes in our French region. We have started the process of moving clients onto a new operational node however our VPSCP is having issues with the migration. We have engaged our vendor that owns the VPSCP software and are awaiting further updates from them for us to resolve this issue.

Fixed

1 year ago

The migration has fully completed and we're back up and running. We'll continue to keep an eye but all appears to be operational at this time.

Watching

1 year ago

The migration is now underway and we think it should be around 2 hours from now before full service is restored at the current transfer times.

Reported

1 year ago

Starting 19:30 BST we will commence migration of the Universeodon.com database server to new infrastructure to enable us to use additional capacity that has been provisioned.

Fixed

1 year ago

Content processing is fully restored and operational.

Identified

1 year ago

We have identified the root cause as a misconfiguration on one of our core routers. We are actively correcting the configuration now and hope to bring the content processing back online shortly.

Reported

1 year ago

We are currently experiencing a major outage on our content processing services. We are looking to restore service ASAP and are actively investigating this issue.

Fixed

1 year ago

Content processing is now fully operational and working as expected.

Watching

1 year ago

We are starting to once again see further performance issues to the database layer of MastodonAppUK which is also resulting in disruption to our ability to process content on Universeodon.com - We are working to mitigate the impacts of this now.

Watching

1 year ago

We have managed to slightly increase our ingress capacity for content processing. We're currently running approx 8 Mins behind live on all our queues with the exception of ingress which is at around 21 Hours behind. We expect this queue to take a couple of hours to fully clear and will continue to monitor.

Watching

1 year ago

We have increased capacity on our database infrastructure however are still hitting bottlenecks. As a result we've scaled back our processing workers to prioritise the default content queue and are currently running a small amount of capacity for the ingress queue to try to catch this up. This means for all queues other than ingress we're currently running around 30 mins behind with ingress currently around 22 hours behind. We will continue to adjust the scaling to ensure the site remains online and operational while we process all this content.

Identified

1 year ago

We have identified a capacity bottleneck on the router that serves part of our database infrastructure, we are scaling this up now and hope this should upgrading some of the bottleneck issues.

Identified

1 year ago

It appears that the content processing has resulted in too much pressure being put on our database infrastructure causing major outages across the site. We are looking to scale back the content processing to restore the sites access.

Watching

1 year ago

We have powered on our legacy content processing server which is starting to work through the backlog, it looks like around midnight on the 29th August 2024 the new content processing services had a major failure resulting in the vast majority of content processing jobs failing to be executed. We currently have a backlog of a little over 1.1 million events which is likely to increase as we process content and additional processing is required. I suspect it'll take a few hours to get things back caught up. We are going to monitor the infrastructure and queues over the coming hours to ensure full recovery.

Reported

1 year ago

Our content processing service has experienced a catastrophic failure resulting in feeds not being updated. We are actively back queueing all of these actions and as soon as we can remediate the issue we will start to catch up on the content that needs processing.

No incidents reported.

Fixed

1 year ago

Issue was previously resolved.

Watching

1 year ago

New infrastructure has been deployed, we will monitor the queue lengths over the next 24 hours to see if we need to deploy further infrastructure to replace our original queue service or if there is tuning we need to do to better make use of the new infrastructure.

Reported

1 year ago

There is currently a major outage on our content processing and queue service. This is due to ongoing maintenance which has not gone to plan. We are currently deploying additional infrastructure to ultimately minimise and mitigate the disruption.

No incidents reported.