TingleFlow Addresses Unplanned Service Outage





Spread the love

July 22, 2026 | TingleFlow | FOR IMMEDIATE RELEASE

TingleFlow experienced an unplanned service outage beginning on July 21, 2026, affecting access to the TingleFlow platform. We understand the importance of providing a reliable and accessible service to our users, and our engineering team has been working continuously to investigate the incident, identify the underlying cause, and restore full functionality.

What Happened:

The outage is related to performance issues within TingleFlowโ€™s internal service discovery infrastructure. During the incident, it began experiencing significant delays with KV write operations, causing extended blocking periods known as contention.

The exact cause of this outage has not yet been fully identified. Our investigation has ruled out several potential causes, including DDoS or cyber attacks, frontend issues, and authentication problems.

Early in the incident, we explored several possible solutions, including restoring the cluster from a healthy snapshot, resetting internal state, adjusting traffic flow, and evaluating hardware performance. While some approaches initially showed improvement, performance degraded again when normal internal service activity resumed.

Further analysis of debug logs, operating system-level metrics, HTOP data, and performance information indicated that the move from 64 CPU Core servers to 128 CPU Core servers during the outage may have contributed to instability. Based on these findings, the engineering team has determined that returning to the previous 64 CPU Core configuration is the best path forward.

Current Status:

TingleFlow remains unavailable while our engineering team prepares the infrastructure transition and continues working toward service restoration.

The team has completed preparation of the replacement hardware environment, reviewing operating system configurations, and performing extensive validation checks. The next step is transitioning the cluster back to 64 CPU Core servers and monitoring system stability.

At this stage, we are unable to provide a confirmed restoration time. Further updates will be shared as progress is made.

Our Response:

Since the outage began, our engineering team has been actively investigating the issue through multiple diagnostic approaches. This has included:

  • Reviewing internal infrastructure behaviour and performance metrics
  • Analysing system logs and operating system-level data
  • Restoring from a previously healthy snapshot
  • Controlling internal traffic using iptables to isolate potential causes
  • Evaluating hardware performance and infrastructure changes
  • Preparing a rollback to the previous server configuration

The team has remained focused on identifying the root cause rather than applying temporary fixes that may not provide long-term stability.

Next Steps:

Our immediate priority is restoring TingleFlow services safely and reliably. The engineering team will continue monitoring performance during and after the infrastructure transition, while investigating the underlying cause of the contention issue.

We will continue providing updates through our status page as new information becomes available.

We sincerely apologise for the disruption this outage has caused. Thank you for your patience and understanding while we work to restore TingleFlow. We will continue to provide updates as they become available to us.

How to stay in the loop

We encourage all users to stay up to date with information so you know when service is restored.

We recommend the status page because you can subscribe to updates using multiple methods.

,


Leave a Reply

Your email address will not be published. Required fields are marked *