# Service deadlock when request volume increases

**URL:** <https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145>\
**Category:** googlegroup\
**Created:** [August 25, 2016, 9:58am UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145 "2016-08-25T09:58:30Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![simon\_harrison](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@simon\_harrison](https://discourse.nameko.io/u/simon_harrison)\
**Post date:** [August 25, 2016, 9:58am UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/1 "2016-08-25T09:58:30Z")

</div>

Hi Devs

We're coming across this situation more and more in existing services and  
new services. The deadlock occurs when one service receives more incoming  
requests than workers are available for it, and as part of each request the  
service entrypoint makes an additional call to another service for data.

E.g.

class FooService:

&nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')

&nbsp;&nbsp;&nbsp;&nbsp;@rpc  
&nbsp;&nbsp;&nbsp;&nbsp;def do\_some\_foo(self):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')

spam = self.spam\_rpc.add\_spam()  
# we never reach here!

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)

If FooService has 10 available workers and 100 requests to `do_some_foo`  
come in quick enough, what appears to happen is that the reply queue set up  
to handle the response from `spam` service fills up (with 10 messages) but  
there are no available workers to handle them. We then also have 10 in  
flight messages in FooService, all un-acked. And here we stay!

Is this behaviour expected/understood? Our assumption is that a new green  
thread should be made available to handle the `spam` responses outside of  
the configured max (10 here).

We're using nameko 2.2.0, RabbitMQ 3.4.3, Python 2.7

We can overcome our immediate issues with design improvements, and indeed  
they will probably prove to be the correct way forward, and we understand  
we can increase the number of workers available, but are concerns are about  
the implementation nameko uses in this situation and how we should design  
services with this in mind.

Thanks for your advice.

---

<div class="post-metadata">

**Author:** ![simon\_harrison](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@simon\_harrison](https://discourse.nameko.io/u/simon_harrison)\
**Post date:** [August 25, 2016, 1:45pm UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/2 "2016-08-25T13:45:39Z")

</div>

\*update\*  
The example I gave was not actually are exact use case.  
All of the entrypoints in question for us are \*event handlers\* which  
subsequently make an RPC call.

&nbsp;&nbsp;&nbsp;&nbsp;class FooService:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;@event\_handler  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;def handle\_foo(self, payload):  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')  
&nbsp;&nbsp;&nbsp;&nbsp;spam = self.spam\_rpc.add\_spam()  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# we never reach here!  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)

Thanks

> **···**
>
> On Thursday, 25 August 2016 10:58:31 UTC+1, simon harrison wrote:
> 
> > Hi Devs
> > 
> > We're coming across this situation more and more in existing services and  
> > new services. The deadlock occurs when one service receives more incoming  
> > requests than workers are available for it, and as part of each request the  
> > service entrypoint makes an additional call to another service for data.
> > 
> > E.g.
> > 
> > class FooService:
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;@rpc  
> > &nbsp;&nbsp;&nbsp;&nbsp;def do\_some\_foo(self):
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')
> > 
> > spam = self.spam\_rpc.add\_spam()  
> > # we never reach here!
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)
> > 
> > If FooService has 10 available workers and 100 requests to `do_some_foo`  
> > come in quick enough, what appears to happen is that the reply queue set up  
> > to handle the response from `spam` service fills up (with 10 messages) but  
> > there are no available workers to handle them. We then also have 10 in  
> > flight messages in FooService, all un-acked. And here we stay!
> > 
> > Is this behaviour expected/understood? Our assumption is that a new green  
> > thread should be made available to handle the `spam` responses outside of  
> > the configured max (10 here).
> > 
> > We're using nameko 2.2.0, RabbitMQ 3.4.3, Python 2.7
> > 
> > We can overcome our immediate issues with design improvements, and indeed  
> > they will probably prove to be the correct way forward, and we understand  
> > we can increase the number of workers available, but are concerns are about  
> > the implementation nameko uses in this situation and how we should design  
> > services with this in mind.
> > 
> > Thanks for your advice.

---

<div class="post-metadata">

**Author:** ![David\_Szotten](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/david_szotten/32/6_2.png) [@David\_Szotten](https://discourse.nameko.io/u/David_Szotten)\
**Post date:** [August 26, 2016, 8:45am UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/3 "2016-08-26T08:45:22Z")

</div>

Hi,

It's been a while since i looked at this in detail, so don't recall all the  
specifics (matt may recall better), but i think there is a known issue due  
to the way consumption from rabbit is set up, and how that interacts with  
the qos settings and max\_workers. this is one of the main things we are  
looking to address by rewriting the amqp internals

d

> **···**
>
> On Thursday, 25 August 2016 14:45:40 UTC+1, simon harrison wrote:
> 
> > \*update\*  
> > The example I gave was not actually are exact use case.  
> > All of the entrypoints in question for us are \*event handlers\* which  
> > subsequently make an RPC call.
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;class FooService:
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;@event\_handler  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;def handle\_foo(self, payload):  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')  
> > &nbsp;&nbsp;&nbsp;&nbsp;spam = self.spam\_rpc.add\_spam()  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# we never reach here!  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)
> > 
> > Thanks
> > 
> > On Thursday, 25 August 2016 10:58:31 UTC+1, simon harrison wrote:
> > 
> > > Hi Devs
> > > 
> > > We're coming across this situation more and more in existing services and  
> > > new services. The deadlock occurs when one service receives more incoming  
> > > requests than workers are available for it, and as part of each request the  
> > > service entrypoint makes an additional call to another service for data.
> > > 
> > > E.g.
> > > 
> > > class FooService:
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;@rpc  
> > > &nbsp;&nbsp;&nbsp;&nbsp;def do\_some\_foo(self):
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')
> > > 
> > > spam = self.spam\_rpc.add\_spam()  
> > > # we never reach here!
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)
> > > 
> > > If FooService has 10 available workers and 100 requests to `do_some_foo`  
> > > come in quick enough, what appears to happen is that the reply queue set up  
> > > to handle the response from `spam` service fills up (with 10 messages) but  
> > > there are no available workers to handle them. We then also have 10 in  
> > > flight messages in FooService, all un-acked. And here we stay!
> > > 
> > > Is this behaviour expected/understood? Our assumption is that a new green  
> > > thread should be made available to handle the `spam` responses outside of  
> > > the configured max (10 here).
> > > 
> > > We're using nameko 2.2.0, RabbitMQ 3.4.3, Python 2.7
> > > 
> > > We can overcome our immediate issues with design improvements, and indeed  
> > > they will probably prove to be the correct way forward, and we understand  
> > > we can increase the number of workers available, but are concerns are about  
> > > the implementation nameko uses in this situation and how we should design  
> > > services with this in mind.
> > > 
> > > Thanks for your advice.

---

<div class="post-metadata">

**Author:** ![mattbennett](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/mattbennett/32/13_2.png) [@mattbennett](https://discourse.nameko.io/u/mattbennett)\
**Post date:** [August 26, 2016, 11:05am UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/4 "2016-08-26T11:05:21Z")

</div>

Indeed, this is another problem caused by the structure of the  
QueueConsumer.

The issue here is that the QoS (aka "prefetch count") is applied once, on  
the channel, rather than for each consumer. The EventHandlers and the RPC  
ReplyListener delegate control of their consumers to the QueueConsumer,  
which means they all share a channel rather than getting their own. The  
prefetch count is pooled for all the consumers on the channel.

The deadlock happens when you exhaust the prefetch count handling events.  
Each concurrent FooService worker holds an unack'd message and reduces the  
prefetch count by one; the message isn't ack'd (and prefetch count  
re-incremented) until the worker completes. But to complete, it must  
consume the RPC reply from the "spam" service. Since the prefetch count is  
already exhausted, they will all wait for somebody else to finish.

There are various ways to work around it, but the ultimate solution is to  
have the AMQP extensions manage their own consumers.

Matt.

> **···**
>
> On Friday, 26 August 2016 09:45:23 UTC+1, David Szotten wrote:
> 
> > Hi,
> > 
> > It's been a while since i looked at this in detail, so don't recall all  
> > the specifics (matt may recall better), but i think there is a known issue  
> > due to the way consumption from rabbit is set up, and how that interacts  
> > with the qos settings and max\_workers. this is one of the main things we  
> > are looking to address by rewriting the amqp internals
> > 
> > d
> > 
> > On Thursday, 25 August 2016 14:45:40 UTC+1, simon harrison wrote:
> > 
> > > \*update\*  
> > > The example I gave was not actually are exact use case.  
> > > All of the entrypoints in question for us are \*event handlers\* which  
> > > subsequently make an RPC call.
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;class FooService:
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;@event\_handler  
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;def handle\_foo(self, payload):  
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')  
> > > &nbsp;&nbsp;&nbsp;&nbsp;spam = self.spam\_rpc.add\_spam()  
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# we never reach here!  
> > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)
> > > 
> > > Thanks
> > > 
> > > On Thursday, 25 August 2016 10:58:31 UTC+1, simon harrison wrote:
> > > 
> > > > Hi Devs
> > > > 
> > > > We're coming across this situation more and more in existing services  
> > > > and new services. The deadlock occurs when one service receives more  
> > > > incoming requests than workers are available for it, and as part of each  
> > > > request the service entrypoint makes an additional call to another service  
> > > > for data.
> > > > 
> > > > E.g.
> > > > 
> > > > class FooService:
> > > > 
> > > > &nbsp;&nbsp;&nbsp;&nbsp;spam\_rpc = RpcProxy('spam')
> > > > 
> > > > &nbsp;&nbsp;&nbsp;&nbsp;@rpc  
> > > > &nbsp;&nbsp;&nbsp;&nbsp;def do\_some\_foo(self):
> > > > 
> > > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('get some spam...')
> > > > 
> > > > spam = self.spam\_rpc.add\_spam()  
> > > > # we never reach here!
> > > > 
> > > > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;logger.info('we have spam: "%s"', spam)
> > > > 
> > > > If FooService has 10 available workers and 100 requests to `do_some_foo`  
> > > > come in quick enough, what appears to happen is that the reply queue set up  
> > > > to handle the response from `spam` service fills up (with 10 messages) but  
> > > > there are no available workers to handle them. We then also have 10 in  
> > > > flight messages in FooService, all un-acked. And here we stay!
> > > > 
> > > > Is this behaviour expected/understood? Our assumption is that a new  
> > > > green thread should be made available to handle the `spam` responses  
> > > > outside of the configured max (10 here).
> > > > 
> > > > We're using nameko 2.2.0, RabbitMQ 3.4.3, Python 2.7
> > > > 
> > > > We can overcome our immediate issues with design improvements, and  
> > > > indeed they will probably prove to be the correct way forward, and we  
> > > > understand we can increase the number of workers available, but are  
> > > > concerns are about the implementation nameko uses in this situation and how  
> > > > we should design services with this in mind.
> > > > 
> > > > Thanks for your advice.

---

<div class="post-metadata">

**Author:** ![Rollo\_Konig\_Brock](https://avatars.discourse-cdn.com/v4/letter/r/58956e/32.png) [@Rollo\_Konig\_Brock](https://discourse.nameko.io/u/Rollo_Konig_Brock)\
**Post date:** [March 23, 2017, 11:11am UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/5 "2017-03-23T11:11:31Z")

</div>

Could you give an example of how to work around it?

---

<div class="post-metadata">

**Author:** ![mattbennett](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/mattbennett/32/13_2.png) [@mattbennett](https://discourse.nameko.io/u/mattbennett)\
**Post date:** [March 23, 2017, 2:32pm UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/6 "2017-03-23T14:32:41Z")

</div>

The workaround I was thinking of was to force the RpcProxy ReplyListener to  
use its own QueueConsumer. You can do this this a series of subclasses:

from nameko.rpc import ReplyListener as NamekoReplyListener  
from nameko.rpc import RpcProxy as RpcProxy  
from nameko.messaging import QueueConsumer as NamekoQueueConsumer

class QueueConsumer(NamekoQueueConsumer):  
&nbsp;&nbsp;&nbsp;&nbsp;@property  
&nbsp;&nbsp;&nbsp;&nbsp;def sharing\_key(self):  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# The `sharing_key` is used to determine whether a new instance of a  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# SharedExtension is needed.  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;# There will always be one instance for each unique sharing\_key  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return "some value"

class ReplyListener(NamekoReplyListener):  
&nbsp;&nbsp;&nbsp;&nbsp;queue\_consumer = QueueConsumer()

class RpcProxy(NamekoRpcProxy):  
&nbsp;&nbsp;&nbsp;&nbsp;rpc\_reply\_listener = ReplyListener()

Nameko 2.5.3 now includes this change  
\<[https://github.com/nameko/nameko/pull/419&gt](https://github.com/nameko/nameko/pull/419&gt); though, which should be  
sufficient to break the deadlock.

> **···**
>
> On Thursday, March 23, 2017 at 11:11:31 AM UTC, rollo.ko...@complyadvantage.com wrote:
> 
> > Could you give an example of how to work around it?

---

<div class="post-metadata">

**Author:** ![Rollo\_Konig\_Brock](https://avatars.discourse-cdn.com/v4/letter/r/58956e/32.png) [@Rollo\_Konig\_Brock](https://discourse.nameko.io/u/Rollo_Konig_Brock)\
**Post date:** [March 23, 2017, 2:57pm UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/7 "2017-03-23T14:57:58Z")

</div>

Adding another prefetched message is not actually enough to break the  
deadlock. In fact the behaviour doesn't change at all.

---

<div class="post-metadata">

**Author:** ![mattbennett](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/mattbennett/32/13_2.png) [@mattbennett](https://discourse.nameko.io/u/mattbennett)\
**Post date:** [March 23, 2017, 3:03pm UTC](https://discourse.nameko.io/t/service-deadlock-when-request-volume-increases/145/8 "2017-03-23T15:03:13Z")

</div>

Can you post an example? I'd be interested to see if it's exactly the same  
behaviour.

Using a dedicated QueueConsumer would be a reliable fix. It will establish  
a new connection for the ReplyListener consumers, which will have their own  
prefetch count.

> **···**
>
> On Thursday, March 23, 2017 at 2:57:58 PM UTC, rollo.ko...@complyadvantage.com wrote:
> 
> > Adding another prefetched message is not actually enough to break the  
> > deadlock. In fact the behaviour doesn't change at all.
