# Serialization cache

**URL:** <https://discourse.nameko.io/t/serialization-cache/139>\
**Category:** googlegroup\
**Created:** [August 14, 2016, 6:51pm UTC](https://discourse.nameko.io/t/serialization-cache/139 "2016-08-14T18:51:38Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![kodonnell](https://avatars.discourse-cdn.com/v4/letter/k/ee7513/32.png) [@kodonnell](https://discourse.nameko.io/u/kodonnell)\
**Post date:** [August 14, 2016, 6:51pm UTC](https://discourse.nameko.io/t/serialization-cache/139/1 "2016-08-14T18:51:38Z")

</div>

It's possibly an edge case, but serialization seems like it can be a  
significant bottleneck with large messages and simple services. So I  
thought an (optional?) serialization and deserialization cache might  
provide some benefit without being too hard to implement (?).

There's also probably some way of going a step further and memoizing -  
allowing the user to specify that this is a deterministic function and if  
that's the case then simply retrieving the result from the cache if it's  
been run before with the same args. (The benefit here is it completely  
skips the serialize -\> send message over network -\> send result over  
network -\> deserialize stuff, which the user can't (easily) avoid, even if  
they memoize their functions directly in their services.)

I'm largely thinking RPC here too.

Thoughts?

---

<div class="post-metadata">

**Author:** ![David\_Szotten](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/david_szotten/32/6_2.png) [@David\_Szotten](https://discourse.nameko.io/u/David_Szotten)\
**Post date:** [August 15, 2016, 3:21pm UTC](https://discourse.nameko.io/t/serialization-cache/139/2 "2016-08-15T15:21:02Z")

</div>

Hi,

Caching of rpc calls is probably best kept outside of the rpc framework  
itself. You could possibly write a wrapper, but this is probably too tied  
to your particular business logic to warrant a blanket solution.

If you are sending large (or otherwise expensive to serialise/deserialise)  
messages you might want to consider using a different serialiser. Iirc,  
kombu has built int (optional) support for msgpack that you can enable  
using the `serializer` config key

Best,  
David

> **···**
>
> On Sunday, 14 August 2016 19:51:38 UTC+1, kodonnell wrote:
> 
> > It's possibly an edge case, but serialization seems like it can be a  
> > significant bottleneck with large messages and simple services. So I  
> > thought an (optional?) serialization and deserialization cache might  
> > provide some benefit without being too hard to implement (?).
> > 
> > There's also probably some way of going a step further and memoizing -  
> > allowing the user to specify that this is a deterministic function and if  
> > that's the case then simply retrieving the result from the cache if it's  
> > been run before with the same args. (The benefit here is it completely  
> > skips the serialize -\> send message over network -\> send result over  
> > network -\> deserialize stuff, which the user can't (easily) avoid, even if  
> > they memoize their functions directly in their services.)
> > 
> > I'm largely thinking RPC here too.
> > 
> > Thoughts?

---

<div class="post-metadata">

**Author:** ![mattbennett](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.nameko.io/mattbennett/32/13_2.png) [@mattbennett](https://discourse.nameko.io/u/mattbennett)\
**Post date:** [August 16, 2016, 6:41am UTC](https://discourse.nameko.io/t/serialization-cache/139/3 "2016-08-16T06:41:49Z")

</div>

I agree with David that caching should be separate from the built-in RPC  
implementation.

Memoizing service methods is a nice pattern that is easily implemented and  
very clean. But as you say, it doesn't avoid having to  
serialize/deserialize on its own.

You could build a cache into the RpcProxy, as you describe, and save even  
going over the network. The RpcProxy is not easy to extend at the moment  
though. It's of limited value if you have lots of different callers too.

At the service end, the best way to implement a cache would be to write a  
custom serializer for kombu. Nameko supports this, and you can see a toy  
implementation in the tests  
\<[https://github.com/onefinestay/nameko/blob/master/test/test\_serialization.py#L183&gt;\](https://github.com/onefinestay/nameko/blob/master/test/test_serialization.py#L183&gt;%5C).  
If you built something like it this it wouldn't be a candidate for  
inclusion in the core library, but I'm sure people would find it useful. It  
would make a great blog post 😉

Matt.

> **···**
>
> On Monday, August 15, 2016 at 11:21:03 PM UTC+8, David Szotten wrote:
> 
> > Hi,
> > 
> > Caching of rpc calls is probably best kept outside of the rpc framework  
> > itself. You could possibly write a wrapper, but this is probably too tied  
> > to your particular business logic to warrant a blanket solution.
> > 
> > If you are sending large (or otherwise expensive to serialise/deserialise)  
> > messages you might want to consider using a different serialiser. Iirc,  
> > kombu has built int (optional) support for msgpack that you can enable  
> > using the `serializer` config key
> > 
> > Best,  
> > David
> > 
> > On Sunday, 14 August 2016 19:51:38 UTC+1, kodonnell wrote:
> > 
> > > It's possibly an edge case, but serialization seems like it can be a  
> > > significant bottleneck with large messages and simple services. So I  
> > > thought an (optional?) serialization and deserialization cache might  
> > > provide some benefit without being too hard to implement (?).
> > > 
> > > There's also probably some way of going a step further and memoizing -  
> > > allowing the user to specify that this is a deterministic function and if  
> > > that's the case then simply retrieving the result from the cache if it's  
> > > been run before with the same args. (The benefit here is it completely  
> > > skips the serialize -\> send message over network -\> send result over  
> > > network -\> deserialize stuff, which the user can't (easily) avoid, even if  
> > > they memoize their functions directly in their services.)
> > > 
> > > I'm largely thinking RPC here too.
> > > 
> > > Thoughts?

---

<div class="post-metadata">

**Author:** ![kodonnell](https://avatars.discourse-cdn.com/v4/letter/k/ee7513/32.png) [@kodonnell](https://discourse.nameko.io/u/kodonnell)\
**Post date:** [August 16, 2016, 6:04pm UTC](https://discourse.nameko.io/t/serialization-cache/139/4 "2016-08-16T18:04:38Z")

</div>

Hi both,

As Matt mentions, it's not really possible outside of the framework. I had  
an idea for doing it in the framework, before realising last night that it  
was very daft. In addition, both of Matt's suggestions are better, so I'll  
have a crack with those. I'm still a tad dinosaur-ish, so don't have a  
blog, but at the least I'll post something on the end of this thread if I  
find a nice implementation = )

Regarding msgpack - yes, I have been using that (along with ujson et. al.).  
It's certainly faster, but the serializing is still the bottleneck by far.  
However, moving to msgpack (or e.g. a custom serialization format) means  
playing with kombu, which fits well with the above suggestion.

Thanks

> **···**
>
> On Tuesday, August 16, 2016 at 6:41:50 PM UTC+12, Matt Yule-Bennett wrote:
> 
> > I agree with David that caching should be separate from the built-in RPC  
> > implementation.
> > 
> > Memoizing service methods is a nice pattern that is easily implemented and  
> > very clean. But as you say, it doesn't avoid having to  
> > serialize/deserialize on its own.
> > 
> > You could build a cache into the RpcProxy, as you describe, and save even  
> > going over the network. The RpcProxy is not easy to extend at the moment  
> > though. It's of limited value if you have lots of different callers too.
> > 
> > At the service end, the best way to implement a cache would be to write a  
> > custom serializer for kombu. Nameko supports this, and you can see a toy  
> > implementation in the tests  
> > \<[https://github.com/onefinestay/nameko/blob/master/test/test\_serialization.py#L183&gt;\](https://github.com/onefinestay/nameko/blob/master/test/test_serialization.py#L183&gt;%5C).  
> > If you built something like it this it wouldn't be a candidate for  
> > inclusion in the core library, but I'm sure people would find it useful. It  
> > would make a great blog post 😉
> > 
> > Matt.
> > 
> > On Monday, August 15, 2016 at 11:21:03 PM UTC+8, David Szotten wrote:
> > 
> > > Hi,
> > > 
> > > Caching of rpc calls is probably best kept outside of the rpc framework  
> > > itself. You could possibly write a wrapper, but this is probably too tied  
> > > to your particular business logic to warrant a blanket solution.
> > > 
> > > If you are sending large (or otherwise expensive to  
> > > serialise/deserialise) messages you might want to consider using a  
> > > different serialiser. Iirc, kombu has built int (optional) support for  
> > > msgpack that you can enable using the `serializer` config key
> > > 
> > > Best,  
> > > David
> > > 
> > > On Sunday, 14 August 2016 19:51:38 UTC+1, kodonnell wrote:
> > > 
> > > > It's possibly an edge case, but serialization seems like it can be a  
> > > > significant bottleneck with large messages and simple services. So I  
> > > > thought an (optional?) serialization and deserialization cache might  
> > > > provide some benefit without being too hard to implement (?).
> > > > 
> > > > There's also probably some way of going a step further and memoizing -  
> > > > allowing the user to specify that this is a deterministic function and if  
> > > > that's the case then simply retrieving the result from the cache if it's  
> > > > been run before with the same args. (The benefit here is it completely  
> > > > skips the serialize -\> send message over network -\> send result over  
> > > > network -\> deserialize stuff, which the user can't (easily) avoid, even if  
> > > > they memoize their functions directly in their services.)
> > > > 
> > > > I'm largely thinking RPC here too.
> > > > 
> > > > Thoughts?

---

<div class="post-metadata">

**Author:** ![kodonnell](https://avatars.discourse-cdn.com/v4/letter/k/ee7513/32.png) [@kodonnell](https://discourse.nameko.io/u/kodonnell)\
**Post date:** [August 16, 2016, 11:20pm UTC](https://discourse.nameko.io/t/serialization-cache/139/5 "2016-08-16T23:20:01Z")

</div>

Hah, it turns out trivial implementations of a generic serialization cache  
don't seem to offer any benefit (they actually slowed my tests down!). The  
main reason being that to cache generic objects, including non-hashable  
ones, you need a way of getting a hash (see how I did it below), and that's  
pretty slow. Ways that could circumvent it:

- if you only use hashable types (probably not useful)  
- the user provides a key which can be used (which gets into the specific  
business case logic mentioned by you both)  
- ??

I guess I should have expected this - creating a hash of an object requires  
similar processes to just serializing. (Indeed, the serialized version  
could be considered a hash.)

Similar considerations will apply to memoizing, I guess.

For completeness, see below for a naive (not production ready!) example of  
a 'caching\_json' kombu serializer (which I registered with setup.py as per  
the docs  
\<[Serialization — Kombu 5.3.4 documentation](http://docs.celeryproject.org/projects/kombu/en/latest/userguide/serialization.html#creating-extensions-using-setuptools-entry-points&gt;%5C)).  
I tested with a simple echo service.

import json  
import sys

loads\_cache = {}  
dumps\_cache = {}

def make\_hashable(value):

&nbsp;&nbsp;&nbsp;&nbsp;if value == None or isinstance(value, (str, int, float)):  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return value  
&nbsp;&nbsp;&nbsp;&nbsp;elif isinstance(value, (tuple, list)):  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return tuple(make\_hashable(v) for v in value)  
&nbsp;&nbsp;&nbsp;&nbsp;elif isinstance(value, (dict)):  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return tuple((k, make\_hashable(v)) for k, v in value.items())  
&nbsp;&nbsp;&nbsp;&nbsp;else:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;raise ValueError("cannot make %s hashable" % str(value)[:100])

def loads(value):

&nbsp;&nbsp;&nbsp;&nbsp;hsh = hash(value)  
&nbsp;&nbsp;&nbsp;&nbsp;cache = globals()['loads\_cache']  
&nbsp;&nbsp;&nbsp;&nbsp;if hsh not in cache:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print('loading')  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;cache[hsh] = json.loads(value)  
&nbsp;&nbsp;&nbsp;&nbsp;return cache[hsh]

def dumps(value):  
&nbsp;&nbsp;&nbsp;&nbsp;hsh = hash(make\_hashable(value))  
&nbsp;&nbsp;&nbsp;&nbsp;cache = globals()['dumps\_cache']  
&nbsp;&nbsp;&nbsp;&nbsp;if hsh not in cache:  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;print('dumping')  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;cache[hsh] = json.dumps(value)  
&nbsp;&nbsp;&nbsp;&nbsp;return cache[hsh]

register\_args = (dumps, loads, 'application/json', 'utf-8')

> **···**
>
> On Wednesday, August 17, 2016 at 6:04:39 AM UTC+12, kodonnell wrote:
> 
> > Hi both,
> > 
> > As Matt mentions, it's not really possible outside of the framework. I had  
> > an idea for doing it in the framework, before realising last night that it  
> > was very daft. In addition, both of Matt's suggestions are better, so I'll  
> > have a crack with those. I'm still a tad dinosaur-ish, so don't have a  
> > blog, but at the least I'll post something on the end of this thread if I  
> > find a nice implementation = )
> > 
> > Regarding msgpack - yes, I have been using that (along with ujson et.  
> > al.). It's certainly faster, but the serializing is still the bottleneck by  
> > far. However, moving to msgpack (or e.g. a custom serialization format)  
> > means playing with kombu, which fits well with the above suggestion.
> > 
> > Thanks
> > 
> > On Tuesday, August 16, 2016 at 6:41:50 PM UTC+12, Matt Yule-Bennett wrote:
> > 
> > > I agree with David that caching should be separate from the built-in RPC  
> > > implementation.
> > > 
> > > Memoizing service methods is a nice pattern that is easily implemented  
> > > and very clean. But as you say, it doesn't avoid having to  
> > > serialize/deserialize on its own.
> > > 
> > > You could build a cache into the RpcProxy, as you describe, and save  
> > > even going over the network. The RpcProxy is not easy to extend at the  
> > > moment though. It's of limited value if you have lots of different callers  
> > > too.
> > > 
> > > At the service end, the best way to implement a cache would be to write a  
> > > custom serializer for kombu. Nameko supports this, and you can see a toy  
> > > implementation in the tests  
> > > \<[https://github.com/onefinestay/nameko/blob/master/test/test\_serialization.py#L183&gt;\](https://github.com/onefinestay/nameko/blob/master/test/test_serialization.py#L183&gt;%5C).  
> > > If you built something like it this it wouldn't be a candidate for  
> > > inclusion in the core library, but I'm sure people would find it useful. It  
> > > would make a great blog post 😉
> > > 
> > > Matt.
> > > 
> > > On Monday, August 15, 2016 at 11:21:03 PM UTC+8, David Szotten wrote:
> > > 
> > > > Hi,
> > > > 
> > > > Caching of rpc calls is probably best kept outside of the rpc framework  
> > > > itself. You could possibly write a wrapper, but this is probably too tied  
> > > > to your particular business logic to warrant a blanket solution.
> > > > 
> > > > If you are sending large (or otherwise expensive to  
> > > > serialise/deserialise) messages you might want to consider using a  
> > > > different serialiser. Iirc, kombu has built int (optional) support for  
> > > > msgpack that you can enable using the `serializer` config key
> > > > 
> > > > Best,  
> > > > David
> > > > 
> > > > On Sunday, 14 August 2016 19:51:38 UTC+1, kodonnell wrote:
> > > > 
> > > > > It's possibly an edge case, but serialization seems like it can be a  
> > > > > significant bottleneck with large messages and simple services. So I  
> > > > > thought an (optional?) serialization and deserialization cache might  
> > > > > provide some benefit without being too hard to implement (?).
> > > > > 
> > > > > There's also probably some way of going a step further and memoizing -  
> > > > > allowing the user to specify that this is a deterministic function and if  
> > > > > that's the case then simply retrieving the result from the cache if it's  
> > > > > been run before with the same args. (The benefit here is it completely  
> > > > > skips the serialize -\> send message over network -\> send result over  
> > > > > network -\> deserialize stuff, which the user can't (easily) avoid, even if  
> > > > > they memoize their functions directly in their services.)
> > > > > 
> > > > > I'm largely thinking RPC here too.
> > > > > 
> > > > > Thoughts?
