R still fits this kind of analysis well, especially once the visualizations matter. The sampling method should be explicit because results can shift depending on whether the data came from a hashtag, a keyword, or replies to one account. Language filtering matters too. I work with AllyHub, and allyhub.com has a related social-data workflow. Did you exclude retweets, or treat them as separate observations?