GraphQL in production - Analyzing public GraphQL APIs #1: Twitch.tv - WunderGraph
Jens Neuse
CEO & Co-Founder at WunderGraph
October 12, 2021·16min read
Last updated on September 9, 2025
State of GraphQL Federation 2026
How are teams governing schema changes, handling production traffic, and measuring Federation success? Share your experience and get early access to the full report. For every valid survey completed, we'll donate $30 to UNICEF.
TL;DR
Twitch hosts an unversioned GraphQL API on a subdomain, sends all requests as HTTP/1.1 POST, and batches roughly 74 operations into about 12 requests with a waterfall exceeding four seconds. Because reads use POST, none of them benefit from Cache-Control headers or ETags, and application-layer batching forces every operation in a batch to wait for the slowest one. Twitch uses Automatic Persisted Queries and blocks introspection, though the schema can still be reconstructed via a known graphql-js suggestions exploit. Better production practice is to send read requests as HTTP GET over HTTP/2 with per-operation caching and ETags, compile operations instead of using APQ, and use HTTP/2 streams instead of WebSockets for realtime updates.
Analyzing public GraphQL APIs is a Series of blog posts to learn from big public GraphQL implementations, starting with Twitch.tv, the popular streaming platform.
We usually assume that GraphQL is just GraphQL. With REST, there's a lot of confusion what it actually is. Build a REST API and the first response you get is that someone says this is not really REST but just JSON over HTTP, etc...
But is this really exclusively a REST thing? Is there really just one way of doing GraphQL?
I've looked at many publicly available GraphQL APIs of companies whose name you're familiar with and analyzed how they "do GraphQL". I quickly realized that everybody does it a bit differently. With this series of posts, I want to extract good and bad patterns from large GraphQL production deployments.
At the end of the series, we'll conclude with a WhitePaper, summarizing all the best practices on how to run GraphQL in production. Make sure to sign up with our WhitePaper early access list. We'll keep you updated on the next post of this series and send you the WhitePaper once it's out.
Analyzing the GraphQL API of Twitch.tv
The first thing you notice is that twitch hosts their GraphQL API on the subdomain https://gql.twitch.tv/gql. Looking at the URL patterns and Headers, it seems that twitch is not versioning their API.
If you look at the Chrome Devtools or similar, you'll notice that for each new "route" on the website, multiple requests are being made to the gql subdomain. In my case, I can count 12 requests on the initial load of the site.
What's interesting is that these requests are being queued sequentially. Starting with the first one at 313ms, then 1.27s, 1.5s, 2.15s, ... , and the last one at 4.33s. One of the promises of GraphQL is to solve the Waterfall problem. However, this only works if all the data required for the website is available in a single GraphQL Operation.
In case of twitch, we've counted 12 requests, but we're not yet at the operation level. Twitch batches requests, but we'll come to that in a minute.
I've noticed another problem with the twitch API. It's using HTTP/1.1 for all requests, not HTTP/2. Why is it a problem? HTTP/2 multiplexes multiple Requests over a single TCP connection, HTTP/1.1 doesn't. You can see this if you look at the timings in Chrome DevTools. Most of the requests can (re-)use an existing TCP Connection, while others initiate a new one. Most of the requests have ~300ms latency while the ones with a connection init and TLS handshake clock in at around 430ms.
Now let's have a closer look at the requests itself. Twitch sends GraphQL Queries using HTTP POST. Their preferred Content-Encoding for Responses is gzip, they don't support brotli.
If you're not logged in, the client sends the Header "Authorization: undefined", which looks like a frontend glitch. Content-Type of the Request is "text/plain" although the payload is JSON.
Some of their requests are single GraphQL requests with a JSON Object. Others are using a batching mechanism, meaning, they send multiple Operations as an Array. The response also comes back as an Array as well, so the client then matches all batched operations to the same response index.
Here's an example of such a batch request:
[
{
"operationName": "ConnectAdIdentityMutation",
"variables": {
"input": {
"targetDeviceID": "2a38ce069ff87bd4"
}
},
"extensions": {
"persistedQuery": {
"version": 1,
"sha256Hash": "aeb02ffde95392868a9da662631090526b891a2972620e6b6393873a39111564"
}
}
},
{
"operationName": "VideoPreviewOverlay",
"variables": {
"login": "dason"
},
"extensions": {
"persistedQuery": {
"version": 1,
"sha256Hash": "3006e77e51b128d838fa4e835723ca4dc9a05c5efd4466c1085215c6e437e65c"
}
}
}
]
Counting all GraphQL Operations for the initial Website load, I get at 74 Operations in total.
Here's a list of all Operations in order of appearance:
- Single 1 (1.2kb Response gzip)
- PlaybackAccessToken_Template
- Batch 1 (5.9kb Response gzip)
- Consent
- Ads_Components_AdManager_User
- Prime_PrimeOffers_CurrentUser
- TopNav_CurrentUser
- PersonalSections
- PersonalSections (different arguments)
- SignupPromptCategory
- ChannelShell
- ChannelVideoLength
- UseLive
- ActiveWatchParty
- UseViewCount
- UseHosting
- DropCurrentSessionContext
- VideoPreviewOverlay
- VideoAdBanner
- ExtensionsOverlay
- MatureGateOverlayBroadcaster
- VideoPlayer_AgeGateOverlayBroadcaster
- CountessData
- VideoPlayer_VideoSourceManager
- StreamTagsTrackingChannel
- ComscoreStreamingQuery
- StreamRefetchManager
- AdRequestHandling
- NielsenContentMetadata
- ExtensionsForChannel
- ExtensionsUIContext_ChannelID
- PlayerTrackingContextQuery
- VideoPlayerStreamMetadata
- Batch 2 (0.7kb Response gzip)
- WatchTrackQuery
- VideoPlayerStatusOverlayChannel
- Batch 3 (20.4 Response gzip)
- ChatRestrictions
- MessageBuffer_Channel
- PollsEnabled
- CommunityPointsRewardRedemptionContext
- ChannelPointsPredictionContext
- ChannelPointsPredictionBadges
- ChannelPointsContext
- ChannelPointsGlobalContext
- ChatRoomState
- Chat_ChannelData
- BitsConfigContext_Global
- BitsConfigContext_Channel
- StreamRefetchManager
- ExtensionsForChannel
- Batch 4 (0.5kb Response gzip)
- RadioCurrentlyPlaying
- Batch 5 (15.7kb Response gzip)
- ChannelPollContext_GetViewablePoll
- AvailableEmotesForChannel
- TrackingManager_RequestInfo
- Prime_PrimeOffers_PrimeOfferIds_Eligibility
- ChatList_Badges
- ChatInput
- VideoPlayerPixelAnalyticsUrls
- VideoAdRequestDecline
- Batch 6 (2kb Response gzip)
- ActiveWatchParty
- UseLive
- RealtimeStreamTagList
- StreamMetadata
- UseLiveBroadcast
- Batch 7 (1.1kb Response gzip)
- ChannelRoot_AboutPanel
- GetHypeTrainExecution
- DropsHighlightService_AvailableDrops
- CrowdChantChannelEligibility
- Batch 8 (1.5kb Response gzip)
- ChannelPage_SubscribeButton_User
- ConnectAdIdentityMutation
- Batch 9 (1.0kb Response gzip)
- RealtimeStreamTagList
- RadioCurrentlyPlaying
- ChannelPage_SubscribeButton_User
- ReportMenuItem
- Batch 10 (1.3kb Response gzip)
- AvailableEmotesForChannel
- EmotePicker_EmotePicker_UserSubscriptionProducts
- Batch 11 (11.7kb Response gzip)
- ChannelLeaderboards
All responses cumulated clock in at 63kb gzipped.
Note that all of these Requests are HTTP POST and therefore don't make any use of Cache-Control Headers. The batch requests use transfer-encoding chunked.
However, on subsequent routes, there seems to be some client-side caching happening. If I change the route to another channel, I can only count 69 GraphQL Operations.
Another observation I can make is that twitch uses APQ, Automatic Persisted Queries. On the first request, the client sends the complete Query to the server. The server then uses the "extends" field on the response object to tell the client the Persisted Operation Hash. Subsequent client requests will then omit the Query payload and instead just send the Hash of the Persisted Operation. This saves bandwidth for subsequent requests.
Looking at the Batch Requests, it seems that the "registration" of Operations happens at build time. So there's no initial registration step. The client only sends the Operation Name as well the Query Hash using the extensions field in the JSON request. (see the example request from above)
Next, I've tried to use Postman to talk to the GraphQL Endpoint.
The first response I've got was a 400, Bad Request.
I've copy-pasted the Client-ID from Chrome Devtools to solve the "problem".
I then wanted to explore their schema. Unfortunately, I wasn't able to use the Introspection Query, it seems to be silently blocked.
However, you could still easily extract the schema from their API using a popular exploit of the graphql-js library.
Finally, I've tried to figure out how their chat works and if they are using GraphQL Subscriptions as well. Switching the Chrome Dev Tools view to "WS" (WebSocket) shows us two WebSocket connections.
One is hosted on the URL wss://pubsub-edge.twitch.tv/v1. It seems to be using versioning, or at least they expect to version this API. Looking at the messages going back and forth between client and server, I can say that the communication protocol is not GraphQL. The information exchanged over this connection is mainly around video playback, server time and view count, so it's keeping the player information in sync.
Discussion
Let's start with the things that surprised me the most.
HTTP 1.1 vs. HTTP2 - GraphQL Request Batching
Batching can be achieved in different ways. One way of batching leverages the HTTP protocol, but batching is also possible in the application layer itself.
Batching has the advantage that it can reduce the number of HTTP requests. In case of twitch, they are batching their 70+ Operations over 12 HTTP requests. Without batching, the Waterfall could be even more extreme. So, it's a very good solution to reduce the number of Requests.
However, batching in the application layer also has its downsides. If you batch 20 Operations into one single Request, you always have to wait for all Operations to resolve before the first byte of the response can be sent to the client. If a single resolver is slow or times out, I assume there are timeouts, all other Operations must wait for the timeout until the responses can be delivered to the client.
Another downside is that batch requests almost always defeat the possibility of HTTP caching. As the API from twitch uses HTTP POST for READ (Query) requests, this option is already gone though.
Additionally, batching can also lead to a slower perceived user experience. A small response can be parsed and processed very quickly by a client. A large response with 20+ kb of gzipped JSON takes longer to parse, leading to longer processing times until the data can be presented in the UI.
So, batching can reduce network latency, but it's not free.
One way of batching makes use of HTTP/2. It's a very elegant way and almost invisible.
HTTP/2 allows browsers to send hundreds of individual HTTP Requests over the same TCP connection. Additionally, the protocol implements Header Compression, which means that client and server can build a dictionary of words in addition to some well known terms to reduce the size of Headers dramatically.
This means, if you're using HTTP/2 for your API, there's no real benefit of "batching at the application layer".
The opposite is actually the case, "batching" over HTTP/2 comes with big advantages over HTTP/1.1 application layer batching.
First, you don't have to wait for all Requests to finish or time out. Each individual request can return a small portion of the required data, which the client can then render immediately.
Second, serving READ Requests over HTTP GET allows for some extra optimizations. You're able to use Cache-Control Headers as well as ETags. Let's discuss these in the next section.
HTTP POST, the wrong way of doing READ requests
Twitch is sending all of their GraphQL Requests over HTTP/1.1 POST. I've investigated the payloads and found out that many of the Requests are loading public data that uses the current channel as a variable. This data seems to be always the same, for all users.
In a high-traffic scenario where millions of users are watching a game, I'd assume that thousands of watchers will continually leave and join the same channel. With HTTP POST and no Cache-Control or ETag Headers, all these Requests will hit the origin server. Depending on the complexity of the backend, this could actually work, e.g. with a REST API and an in memory database.
However, these POST Requests hit the origin server which then executes the persisted GraphQL Operations. This can only work with thousands of servers, combined with a well-defined Resolver architecture using the Data-Loader pattern and application-side caching, e.g. using Redis.
I've looked into the Response timings, and they are coming back quite fast! So, the twitch engineers must have done a few things quite well to handle this kind of load with such a low latency.
Let's discuss some considerations and suggestions on optimizing GraphQL in production environments.
Suggestions
READ Requests should always use HTTP GET over HTTP/2
READ Requests or GraphQL Queries should always use HTTP GET Requests over HTTP/2, allowing for better performance, caching, and reduced bandwidth utilization.
APQ < Compiled Operations
Instead of relying on Automatic Persisted Queries, it may be beneficial to compile Operations to increase performance.
Subscriptions over HTTP/2 Streams
Utilizing HTTP/2 Streams instead of WebSockets can lead to improvements in handling realtime updates efficiently.
Conclusion
This series aims to shed light on the various approaches to building GraphQL APIs, advocating for a standard approach that maximizes performance and security. We encourage API developers to adopt best practices to streamline operations.