I am trying to understand how to use distributed arrays or spmd code in order to optimize my computation. Specifically, I am trying to implement this in the context of an iterative solver. Essentially, within my algorithm code, I have just converted all of my input and starting vectors to distributed arrays. For example, the right hand side vector b and input matrix A become
A = distributed(A);
b = distributed(b);
The problem is that each iteration of the solver is now taking on the order of 100 times longer per iteration (and thus per solve) compared to when I did not distribute the data.
I am wondering if there is some threshhold at which the distributed memory wins out over not distributing, perhaps if the problem size gets large enough. My initial tests were on relatively small sparse matrices, around 3000×3000. But I have plenty more matrices of larger size to test on, up to around 10E6x10E6.
I also did some rudimentary tests just computing dot products of distributed random vectors. The time to compute a simple dot product where the vector was stored across distributed memory was consistently on the order of 100 times longer to compute compared to not distributing the vector.
rng(‘default’);
v = rand(1,2000000); % Test vector across different sizes
tic
vv_true = v*v’ % 1) compute dot product without distributing v
toc
v = distributed(v);
spmd
v % Verify the dot product distributed across workers
end
tic
vv_dist = v*v’ % 2) compute dot product after distributing v
toc
This simple test is important to me because the iterative solver I am working on uses a substantially smaller number of dot products compared to traditional iterative methods like BiCGStab. In my tests, I really would like to show the speedup from the algorithm structure when the problem is done in the context of distributed memory, as I expect it would be in practice if the solver were used to solve really large data sets.
So, it is in one sense reassuring to see that computing dot products does take longer in a distributed memory context compared to a non-distributed (is there another term?) context. But, in another sense it is incredibly impractical to solve these intensive and time consuming problems if it is actually going to take 100 times longer. I simply don’t even have the compute time on the cluster I work with to turn a 6-hour problem into a 600-hour one.
My question is three-fold:
Am I using distributed memory correctly and in the most efficient way in the context of an iterative solver?
At what point does the overhead cost of transferring the results between workers become smaller than the benefit of having multiple workers do the work?
In the simple dot product code above, should it really take ~100 times longer to compute the dot product on a distributed data framework compared to a non-distributed one?
Thank you!I am trying to understand how to use distributed arrays or spmd code in order to optimize my computation. Specifically, I am trying to implement this in the context of an iterative solver. Essentially, within my algorithm code, I have just converted all of my input and starting vectors to distributed arrays. For example, the right hand side vector b and input matrix A become
A = distributed(A);
b = distributed(b);
The problem is that each iteration of the solver is now taking on the order of 100 times longer per iteration (and thus per solve) compared to when I did not distribute the data.
I am wondering if there is some threshhold at which the distributed memory wins out over not distributing, perhaps if the problem size gets large enough. My initial tests were on relatively small sparse matrices, around 3000×3000. But I have plenty more matrices of larger size to test on, up to around 10E6x10E6.
I also did some rudimentary tests just computing dot products of distributed random vectors. The time to compute a simple dot product where the vector was stored across distributed memory was consistently on the order of 100 times longer to compute compared to not distributing the vector.
rng(‘default’);
v = rand(1,2000000); % Test vector across different sizes
tic
vv_true = v*v’ % 1) compute dot product without distributing v
toc
v = distributed(v);
spmd
v % Verify the dot product distributed across workers
end
tic
vv_dist = v*v’ % 2) compute dot product after distributing v
toc
This simple test is important to me because the iterative solver I am working on uses a substantially smaller number of dot products compared to traditional iterative methods like BiCGStab. In my tests, I really would like to show the speedup from the algorithm structure when the problem is done in the context of distributed memory, as I expect it would be in practice if the solver were used to solve really large data sets.
So, it is in one sense reassuring to see that computing dot products does take longer in a distributed memory context compared to a non-distributed (is there another term?) context. But, in another sense it is incredibly impractical to solve these intensive and time consuming problems if it is actually going to take 100 times longer. I simply don’t even have the compute time on the cluster I work with to turn a 6-hour problem into a 600-hour one.
My question is three-fold:
Am I using distributed memory correctly and in the most efficient way in the context of an iterative solver?
At what point does the overhead cost of transferring the results between workers become smaller than the benefit of having multiple workers do the work?
In the simple dot product code above, should it really take ~100 times longer to compute the dot product on a distributed data framework compared to a non-distributed one?
Thank you! I am trying to understand how to use distributed arrays or spmd code in order to optimize my computation. Specifically, I am trying to implement this in the context of an iterative solver. Essentially, within my algorithm code, I have just converted all of my input and starting vectors to distributed arrays. For example, the right hand side vector b and input matrix A become
A = distributed(A);
b = distributed(b);
The problem is that each iteration of the solver is now taking on the order of 100 times longer per iteration (and thus per solve) compared to when I did not distribute the data.
I am wondering if there is some threshhold at which the distributed memory wins out over not distributing, perhaps if the problem size gets large enough. My initial tests were on relatively small sparse matrices, around 3000×3000. But I have plenty more matrices of larger size to test on, up to around 10E6x10E6.
I also did some rudimentary tests just computing dot products of distributed random vectors. The time to compute a simple dot product where the vector was stored across distributed memory was consistently on the order of 100 times longer to compute compared to not distributing the vector.
rng(‘default’);
v = rand(1,2000000); % Test vector across different sizes
tic
vv_true = v*v’ % 1) compute dot product without distributing v
toc
v = distributed(v);
spmd
v % Verify the dot product distributed across workers
end
tic
vv_dist = v*v’ % 2) compute dot product after distributing v
toc
This simple test is important to me because the iterative solver I am working on uses a substantially smaller number of dot products compared to traditional iterative methods like BiCGStab. In my tests, I really would like to show the speedup from the algorithm structure when the problem is done in the context of distributed memory, as I expect it would be in practice if the solver were used to solve really large data sets.
So, it is in one sense reassuring to see that computing dot products does take longer in a distributed memory context compared to a non-distributed (is there another term?) context. But, in another sense it is incredibly impractical to solve these intensive and time consuming problems if it is actually going to take 100 times longer. I simply don’t even have the compute time on the cluster I work with to turn a 6-hour problem into a 600-hour one.
My question is three-fold:
Am I using distributed memory correctly and in the most efficient way in the context of an iterative solver?
At what point does the overhead cost of transferring the results between workers become smaller than the benefit of having multiple workers do the work?
In the simple dot product code above, should it really take ~100 times longer to compute the dot product on a distributed data framework compared to a non-distributed one?
Thank you! iterative solvers, distributed memory, parallel computing MATLAB Answers — New Questions
