With a model like #cliptagger from $GRASS, models can: - Use the JSON output and "watch" videos, and tell you what is happening and when - Because the output is semantic, they can be used as inputs for reasoning tasks, training video models and even build datasets for robotics.
you can try cliptagger for yourself! upload any image or video frame and run. generation takes a few minutes right now but this'll be blazing fast once we swap in @inference_net on the backend. fun fact, this powers our video search.
Show original
8.66K
34
The content on this page is provided by third parties. Unless otherwise stated, OKX is not the author of the cited article(s) and does not claim any copyright in the materials. The content is provided for informational purposes only and does not represent the views of OKX. It is not intended to be an endorsement of any kind and should not be considered investment advice or a solicitation to buy or sell digital assets. To the extent generative AI is utilized to provide summaries or other information, such AI generated content may be inaccurate or inconsistent. Please read the linked article for more details and information. OKX is not responsible for content hosted on third party sites. Digital asset holdings, including stablecoins and NFTs, involve a high degree of risk and can fluctuate greatly. You should carefully consider whether trading or holding digital assets is suitable for you in light of your financial condition.