Repository navigation
object 'results' not found #334
Description
Activity
Have a look at this issue 284, I have just been running into this myself and it seems using the option setAutoDeleteJob(FALSE) and .options.azure = list(enableCloudCombine = FALSE) will solve your issue. the link has more details but bassically you merge it yourself by reading from the blob storage directly.
Pullarg I tried
my_results <- foreach(t = 1:3, .options.azure = list(enableCloudCombine = FALSE, autoDeleteJob = FALSE)) %dopar% { # object 'results' not found, which you can see in my original post. That should behave the same way as usingsetAutoDeleteJob(FALSE). That being said, I tested several variations withsetAutoDeleteJob(FALSE)anyway (code below). All resulted in the same error (also shown below).Error message
============================================================================== Id: job20181129155730 chunkSize: 1 enableCloudCombine: FALSE errorHandling: stop wait: TRUE autoDeleteJob: FALSE ============================================================================== Submitting tasks (3/3) Waiting for tasks to complete. . . | Progress: 100.00% (3/3) | Running: 0 | Queued: 0 | Completed: 3 | Failed: 0 | Tasks have completed. Error in e$fun(obj, substitute(ex), parent.frame(), e$data) : object 'results' not found Called from: e$fun(obj, substitute(ex), parent.frame(), e$data)Sample code
library(doAzureParallel) setVerbose(TRUE) setAutoDeleteJob(FALSE) setCredentials(file.path(getwd(), "credentials.json")) cluster <- makeCluster(file.path(getwd(), "cluster.json"), fullName=TRUE) registerDoAzureParallel(cluster) getDoParWorkers() # my_results <- foreach(t = 1:3, .options.azure = list(enableCloudCombine = FALSE, autoDeleteJob = FALSE)) %dopar% { # object 'results' not found # my_results <- foreach(t = 1:3, .options.azure = list(enableCloudCombine = FALSE)) %dopar% { # object 'results' not found # my_results <- foreach(t = 1:3, .combine = 'rbind', .options.azure = list(enableCloudCombine = FALSE)) %dopar% { # object 'results' not found my_results <- foreach(t = 1:3, .combine = 'rbind', .options.azure = list(enableCloudCombine = FALSE, autoDeleteJob = FALSE)) %dopar% { # object 'results' not found my_results_df <- data.frame("x" = runif(2), "trial" = replicate(2, t)) my_results_list <- runif(3) return(my_results_df) }If you remove the
enableCloudCombineflag, you will get your results. The object 'result not found' occurs because no file is found on Azure Storage that contains the merged result (RDS file that contains all the tasks sinceenableCloudCombineis set to disable). I will add better error handling for this case.Below: This example works
my_results <- foreach(t = 1:3, .combine = 'rbind') %dopar% { my_results_df <- data.frame("x" = runif(2), "trial" = replicate(2, t)) my_results_list <- runif(3) return(my_results_df) } my_results
Brian Hoang (@brnleehng) Yep, that's what I've been doing.
I suppose I don't understand the use case for
enableCloudCombine = FALSE. How should we be using this option?The documentation doesn't have any clear examples besides what's mentioned here. Looking at that example, I feel like that would also trigger this error.
Brian Hoang (@brnleehng) I need to return a list rather than bind rows. Is there anyway to skip the merge step at all, as i am getting the same failure. I would like to just perform this on the head( non cloud side) by reading from the storage account directly.
Update:
Dumb question just needed to turn the result back into a list. , after having the rbind result i can convert from a data.frame to a list by using this statement split(rbind.df, seq(nrow(rbind.df)))The case for
enableCloudCombine = FALSEis to avoid merging all your resources onto one VM while the other VMs are in idle (Unless you are using autoscale). There are cases when your tasks are producing many/large files that the merge task can run out of memory causing your job to fail.Hi Pullarg,
You should usegetJobResultfunction to download all the results locally and it will manually merge it as a list.> getJobResult("job20181205211937") Getting job results... enableCloudCombine is set to FALSE, we will merge job result locally [[1]] [1] 2 [[2]] [1] 3 [[3]] [1] 4
Thanks,
BrianI'm getting the same error even with
enableCloudCombine = FALSE. In my code, I am not returning any results from the %dopar% block. Instead, I am just writing my result dataframe to disk. My code runs correctly but the error still appears. Is there a way to avoid this error when the code intentionally does not return a result?Can you add NULL at the end of the %dopar% block?
I'm looking into fixing enableCloudCombine path = false.
Before submitting a bug please check the following:
sessionInfo()Updates
EDIT (11/29/2018) - Added additional examples, corrected a typo and improved formatting.
Description
Can someone please explain why I get the error
object 'results' not found? Full code below.Ideally, I need to return two objects from inside the loop. One is a data frame that gets row binded, and the other is a list that needs to become a list of lists. In the example below, I'm only returning the data frame (will work on adding the list as additional output once this issue is resolved).
Instruction to repro the problem if applicable
Example 1
Example 2 featuring superfluous use of
setAutoDeleteJob(FALSE)Output from
sessionInfo()Output from error