ATLAS Media Group Status Page

Some systems are experiencing issues.

ATLAS Media Group Status Page

The latest status updates from the team. Updates provided on known and ongoing incidents.

Planned maintenance

  • We need to perform routine maintenance to our physical hardware located within the Redditch Data center. As part of this we will need to take all services offline for a short amount of time while we apply system and security updates to our shared storage pool and restart the underlying host. We will also be performing patching and upgrades to our application hosts during this time and will be failing over our applications across the nodes in the environment.

    We expect the maintenance itself to take approx 3 hours with a short outage of up to 30 mins throughout this while we take the environment offline for full system restarts and hardware improvements.

    Affected Components: VPS Control Panel, Universeodon Database, Universeodon Queue & Content Processing Service, Universeodon Website, Universeodon Relay, Billing Panel, Blog, MastodonApp.UK Website, MastodonAppUK Database Service and MastodonAppUK Queue & Content Processing Service

Past incidents

No incidents reported.

MastodonAppUK - Redis Migration

6 months ago
Complete
Affected Components: MastodonApp.UK Website, MastodonAppUK Advanced / Full Text Search, MastodonAppUK E-Mail Service and MastodonAppUK Queue & Content Processing Service
6 months ago

We have completed the Redis migration and all services are now fully restored and content processing queues are fully caught up.

Affected Components: Universeodon Queue & Content Processing Service

Fixed

6 months ago

Queues have now fully cleared and we're operating in real time with the wider network of servers.

Watching

6 months ago

Our ingest pipeline is down to running around 8 hours behind live at this time. All other queues appear to be operating as expected.

We are seeing a high error rate when interacting with our object storage endpoint especially when generating preview images for links so a large number of jobs are likely going to need to be manually re-ran.

Watching

6 months ago

We have restarted our content processing services which appear to have hung without any errors being recorded. We will continue to monitor the queues while they reduce. As of this update all queues with the exception of ingress are now running in real time and ingress is tracking approx 13 hours behind.

Reported

6 months ago

We are investigating reports of the queue service and content processing service being unresponsive.

No incidents reported.

No incidents reported.

No incidents reported.

Redditch DC Maintenance - Risk to service

6 months ago
Complete
6 months ago

Maintenance has been completed.

6 months ago

We've completed the install on APP-1 at the Redditch DC, we're in the process of applying OS updates now to the hypervisor and will commence a failover from APP-2 back to APP-1 ready to take APP-2 out of service for the same upgrade.

6 months ago

We have completed the first step of the works which has increased our disk capacity for our SSD backed storage at the site and confirmed all is now operational. We will shortly start work to upgrade the networking within our app servers on-site which may involve some sort outages while capacity is moved to other available nodes.