There were not random queries. In fact, they were picked because we felt they were likely to trigger a knowledge panel.
For a good believable test your sampling methodology is very important. In fact, you want to have a random sample that is free of biases instead of somebody cherry picking queries. Perhaps you may do random weighted sample of queries with weights = frequency which represents usage pattern more closely. In any case, it's very important to describe your sampling methodology or otherwise this kind of testing has little value.
If you want to do this kind of test with more discipline, you would give out random phones to randomly selected N people who have never used any of these services before. Then log each query they do for a week or two. Afterwards you can run same queries through all 3 services, generate results and do human judgement on mechanical turk about which one is the best. The science of measuring these stuff is complex and I've omitted many complexities here, for example, you want to do multiple judgement for same query, you need someway to measure judgement quality itself, you need great judgement guidelines that covers edge cases etc. Usually in organizations that work with big data, you will find competent measurement team building tools for these kind of measurements with years of investment.
I don't think it's possible, people on different platforms issue different queries partially due to marketing and insiticts on what works and what doesn't.
Also commands that already sort of work lead to more queries issues lead to better quality.
For a good believable test your sampling methodology is very important. In fact, you want to have a random sample that is free of biases instead of somebody cherry picking queries. Perhaps you may do random weighted sample of queries with weights = frequency which represents usage pattern more closely. In any case, it's very important to describe your sampling methodology or otherwise this kind of testing has little value.