Efficiently choosing a random line from a text file with uniform probability in C?

Question 1

Select a random character from the file (via rand and seek as you noted). Now, instead of finding the associated newline, since that is biased as you noted, I would apply the following algorithm:


Is the character a newline character?
   yes - use the preceeding line
   no  - try again

I can't see how this could give anything but a uniform distribution of lines. The efficiency depends on the average length of a line. If your file has relatively short lines, this could be workable, though if the file really can't be precached even by the OS, you might pay a heavy price in physical disk seeks.

Question 2

Solution was found, which works surprisingly well. Documenting here for myself and others.

This example code does around 80,000 draws per second in practice, with a mean line length that matches that of the file to 4 significant digits on most runs. In contrast, I get around 250 draws per second using the method from the cross referenced question.

Essentially what it does is sample a random place in the file, and then discard it and draw again with probability inversely proportionate to the line length. This cancels out the bias for longer words. On average, the method makes a number of draws equal to the average line length in the file before accepting one.

Some notable drawbacks:

Files with longer line lengths will produce more rejections per draw, making this much slower.

Files with longer line lengths require a larger constant than 50 in the rdraw function, which appears to mean much longer seek times in practice if line lengths exhibit high variance. For instance, setting it to BUFSIZ on one file I tested with reduced draw speeds to around 10000 draws per second. Still much faster than counting lines in the file though.

int rdraw(FILE* where, char *storage, size_t bytes){
    int offset = (int)(bytes*drand48());
    int initial_seek = offset>50?offset-50:0;
    fseek(where, initial_seek, SEEK_SET);
    int chars_read = 0;
    while(chars_read + initial_seek < offset){
            fgets(storage,50,where);
            chars_read += strlen(storage);
    }
    return strlen(storage);
}

int main(){
    srand48(time(NULL));
    struct stat blah;
    stat("/usr/share/dict/words", &blah);
    FILE *where = fopen("/usr/share/dict/words", "r");
    off_t bytes = blah.st_size;
    char b[BUFSIZ+1];

    int i;
    for(i=0;i<1000000; i++){
            while(drand48() > 1.0/(rdraw(where, b, bytes)));
    }

}

Question 3

If the file only changes in the end (more lines are added) you can create an algorithm with uniform probability:

Preparation: Create an index file that contains the offset for each n:th line. Use a fixed-width format so that the position can be used to determine which record you have.

Open the index file and read the last record. Use ftell to determine the record number.
Open the big file and fseek to the offset obtained in step 1.
Read the big file to the end, counting the number of newlines. You now have the total number of lines in the big file.
Generate a random number up to the number of lines obtained in step 3.
fseek to, and read, the appropriate record in the index file.
fseek to the appropriate offset in the large file. Skip the remainder of newlines.
Read the line!

Example

Let's assume we chose n=100 and that the large file contains 367 lines.

Index file:

00000000,00004753,00009420,00016303

The index file has 4 records, so the large file containsat least 300 records (100* (4-1)). Last offset is 16303.
Open the large file and fseek to 16303.
Count the remaining number of lines (67).
Generata a random number in the range [0-366]. Let's say we got 112.
112/100 = 1 with 12 as remainder. Read the index file record with offset 1. We get the result 4753.
fseek to 4753 in the large file and then skip 11 (12-1) lines.
Read the 12th line.

Voila!

Edit:

I saw the comment on the target file changing. If the target file changes only rarely, then this may still be a viable approach. You would need to create an new index file before switching target file. You may also want to update the index file when the target file has grown more than n rows.