[SPARK-27629][PySpark] Prevent Unpickler from intervening each unpickling #24521

viirya · 2019-05-03T16:06:56Z

What changes were proposed in this pull request?

In SPARK-27612, one correctness issue was reported. When protocol 4 is used to pickle Python objects, we found that unpickled objects were wrong. A temporary fix was proposed by not using highest protocol.

It was found that Opcodes.MEMOIZE was appeared in the opcodes in protocol 4. It is suspect to this issue.

A deeper dive found that Opcodes.MEMOIZE stores objects into internal map of Unpickler object. We use single Unpickler object to unpickle serialized Python bytes. Stored objects intervenes next round of unpickling, if the map is not cleared.

We has two options:

Continues to reuse Unpickler, but calls its close after each unpickling.
Not to reuse Unpickler and create new Unpickler object in each unpickling.

This patch takes option 1.

How was this patch tested?

Passing the test added in SPARK-27612 (#24519).

viirya · 2019-05-03T16:07:11Z

cc @HyukjinKwon @BryanCutler

SparkQA · 2019-05-03T17:44:12Z

Test build #105107 has finished for PR 24521 at commit d7312fb.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

dongjoon-hyun · 2019-05-03T17:59:16Z

Retest this please.

BryanCutler

Thanks for digging into this @viirya , I'm still trying to understand the reason why previous data in the memoize map creates the output we see. Ultimately, this does look like a bug in Pyrolite and this workaround might be fine. Would you mind adding a reference to the JIRA in your comment and in the test added in #24519?

BryanCutler · 2019-05-03T17:59:28Z

core/src/main/scala/org/apache/spark/api/python/SerDeUtil.scala

@@ -186,6 +186,9 @@ private[spark] object SerDeUtil extends Logging {
      val unpickle = new Unpickler
      iter.flatMap { row =>
        val obj = unpickle.loads(row)
+        // `Opcodes.MEMOIZE` of Protocol 4 (Python 3.4+) will store objects in internal map
+        // of `Unpickler`. This map is cleared when calling `Unpickler.close()`.
+        unpickle.close()


It looks like close() clears the memoized map, the UnpickleStack, and the input stream. Since the input stream is a ByteArrayStream that's a no-op and is fine. Since each row is independent of the others, I don't see any reason why the other 2 would store anything necessary, so I believe that will be fine.

I don't know anything about this, but it looks odd to close the object repeatedly. It may not cause a problem now. What's the downside to using a new object for each row, just performance?

Yea, I missed Since each row is independent of the others - so I thought we should manually control memo. I had to look deeper :D. Looks fine. Yes, I guess just it needs an extra clear call for each row.

To use use a new object for each row is also ok. I didn't do it just for (possible) performance concern.

SparkQA · 2019-05-03T20:36:50Z

Test build #105112 has finished for PR 24521 at commit d7312fb.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2019-05-03T23:30:42Z

Nice try. My impression was that it somehow wrongly reuses that memo (and can't close the input stream - I overlooked) so was thinking we need a manual control after upgrading to 4.17 (so that we can access to memo) but looking into this again the fix looks correct.

HyukjinKwon

@viirya, BTW, I believe there are some more places to fix, for instance, BatchEvalPythonExec. Can you double check those places too while we're here?

viirya · 2019-05-04T01:35:03Z

Ultimately, this does look like a bug in Pyrolite and this workaround might be fine. Would you mind adding a reference to the JIRA in your comment and in the test added in #24519?

I'd like to file an issue to Pyrolite. From my view, memo map should be cleared when op code STOP is hit, because each loading should be independent, so the memo map shouldn't be reuse without clearing.

I added few comments to the test and the JIRA.

viirya · 2019-05-04T01:41:30Z

BTW, I believe there are some more places to fix, for instance, BatchEvalPythonExec. Can you double check those places too while we're here?

Checked other places and fixed that. Thanks for reminding that.

SparkQA · 2019-05-04T03:04:51Z

Test build #105117 has finished for PR 24521 at commit 04a2e04.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2019-05-04T04:00:46Z

Test build #105118 has finished for PR 24521 at commit 053e6a5.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon

LGTM nice followup

HyukjinKwon · 2019-05-04T04:20:24Z

Merged to master.

gatorsmile · 2019-05-08T03:27:31Z

@viirya Could you please help measure the perf?

HyukjinKwon · 2019-05-08T03:31:06Z

Shouldn't be a big diff and I don't think there's other ways to fix this issue even if there's perf diff.
But yes, let's quickly run a simple job and check before/after to clarify the doubt. This was pointed out by @srowen too.

viirya · 2019-05-08T03:46:37Z

I'm ok. I also think it shouldn't be a big diff. I will post the perf diff.

viirya · 2019-05-08T08:47:56Z

I ran a pretty simple job and measure time before and after this PR in 5 runs:

>>> import time
>>> start = time.time()
>>> df = spark.createDataFrame([[1, 2, 3, 4]] * 100000, "array<integer>")
>>> collected = df.collect()
>>> end = time.time()
>>> print(end - start)

Before:
0.6398153305053711
0.6232590675354004
0.603309154510498
0.5923750400543213
0.5802857875823975

After:
0.6308741569519043
0.6202919483184814
0.5926861763000488
0.603518009185791
0.5818917751312256

…ling In SPARK-27612, one correctness issue was reported. When protocol 4 is used to pickle Python objects, we found that unpickled objects were wrong. A temporary fix was proposed by not using highest protocol. It was found that Opcodes.MEMOIZE was appeared in the opcodes in protocol 4. It is suspect to this issue. A deeper dive found that Opcodes.MEMOIZE stores objects into internal map of Unpickler object. We use single Unpickler object to unpickle serialized Python bytes. Stored objects intervenes next round of unpickling, if the map is not cleared. We has two options: 1. Continues to reuse Unpickler, but calls its close after each unpickling. 2. Not to reuse Unpickler and create new Unpickler object in each unpickling. This patch takes option 1. Passing the test added in SPARK-27612 (apache#24519). Closes apache#24521 from viirya/SPARK-27629. Authored-by: Liang-Chi Hsieh <viirya@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>

Prevent Unpickler from intervening each unpickling by calling close.

d7312fb

BryanCutler reviewed May 3, 2019

View reviewed changes

HyukjinKwon approved these changes May 3, 2019

View reviewed changes

HyukjinKwon reviewed May 3, 2019

View reviewed changes

Add comment.

04a2e04

Call close on Unpickler after loading.

053e6a5

HyukjinKwon approved these changes May 4, 2019

View reviewed changes

HyukjinKwon closed this in d9bcacf May 4, 2019

rshkv mentioned this pull request May 23, 2020

Misc PyArrow fixes palantir/spark#684

Merged

9 tasks

viirya deleted the SPARK-27629 branch December 27, 2023 18:22

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-27629][PySpark] Prevent Unpickler from intervening each unpickling #24521

[SPARK-27629][PySpark] Prevent Unpickler from intervening each unpickling #24521

viirya commented May 3, 2019 •

edited by HyukjinKwon

Loading

viirya commented May 3, 2019

SparkQA commented May 3, 2019

dongjoon-hyun commented May 3, 2019

BryanCutler left a comment

BryanCutler May 3, 2019

srowen May 3, 2019

HyukjinKwon May 3, 2019

viirya May 4, 2019

SparkQA commented May 3, 2019

HyukjinKwon commented May 3, 2019 •

edited

Loading

HyukjinKwon left a comment

viirya commented May 4, 2019

viirya commented May 4, 2019

SparkQA commented May 4, 2019

SparkQA commented May 4, 2019

HyukjinKwon left a comment

HyukjinKwon commented May 4, 2019

gatorsmile commented May 8, 2019

HyukjinKwon commented May 8, 2019

viirya commented May 8, 2019

viirya commented May 8, 2019

[SPARK-27629][PySpark] Prevent Unpickler from intervening each unpickling #24521

[SPARK-27629][PySpark] Prevent Unpickler from intervening each unpickling #24521

Conversation

viirya commented May 3, 2019 • edited by HyukjinKwon Loading

What changes were proposed in this pull request?

How was this patch tested?

viirya commented May 3, 2019

SparkQA commented May 3, 2019

dongjoon-hyun commented May 3, 2019

BryanCutler left a comment

Choose a reason for hiding this comment

BryanCutler May 3, 2019

Choose a reason for hiding this comment

srowen May 3, 2019

Choose a reason for hiding this comment

HyukjinKwon May 3, 2019

Choose a reason for hiding this comment

viirya May 4, 2019

Choose a reason for hiding this comment

SparkQA commented May 3, 2019

HyukjinKwon commented May 3, 2019 • edited Loading

HyukjinKwon left a comment

Choose a reason for hiding this comment

viirya commented May 4, 2019

viirya commented May 4, 2019

SparkQA commented May 4, 2019

SparkQA commented May 4, 2019

HyukjinKwon left a comment

Choose a reason for hiding this comment

HyukjinKwon commented May 4, 2019

gatorsmile commented May 8, 2019

HyukjinKwon commented May 8, 2019

viirya commented May 8, 2019

viirya commented May 8, 2019

viirya commented May 3, 2019 •

edited by HyukjinKwon

Loading

HyukjinKwon commented May 3, 2019 •

edited

Loading